About the hub
WorkBench.ai ranks AI models on concrete business tasks, one business function at a time, and publishes everything needed to redo the sums.
What public leaderboards do not measure
Public AI leaderboards measure academic exams. They tell you which model reasons best on that ground, and that is worth knowing. It is not what a company needs to know to choose a tool.
An SME buying a tool is asking other questions:
- how many of its invoices will go through without correction;
- how much each one will cost;
- how often the model will invent a figure that is not on the document.
No exam score answers them. The hub is built to answer them — for invoices first, then for the company's other business functions.
What we build
Concrete business tasks, grouped by business function: finance, accounting, human resources, legal, sales, marketing, customer service, procurement, IT, management. Each benchmark starts from a question an executive would ask — how many of your invoices will go through without a human fixing them? — not from some abstract capability of the model.
Four measures, published side by side and never melted into a single score:
- No review needed: the share of cases handled with nothing for a human to fix.
- Accuracy: the points earned across all graded items, weighted by how critical they are.
- Hallucinations: items invented where the source contains nothing.
- Cost and time: per test case, at public prices on the day of the test.
A model can come first on accuracy and be unusable because it makes things up. A single score would hide exactly that. The site's one overall figure, the Business Index, aggregates accuracy alone, with every business function weighing the same.
Three verdicts rather than two. A field absent from the document is part of the test: an invoice under the VAT exemption has no VAT rate, and answering “nothing” is then the right answer. So a field is correct, wrong or missing, or hallucinated — the model produced a value where the document contains none. An invented figure costs more than a missing one, because it gets through review.
Everything can be checked
The prompt sent, the documents submitted, each model's raw answers, the scoring grid and the scoring code are published in the repository. Every measured figure can be recomputed from there.
A run is a timestamped, immutable folder. Adding a model creates a new run; earlier ones remain available. The exact version of the model that answered is recorded at run time, never typed in by hand: six months from now, the same commercial name may no longer point to the same model.
The steps are kept separate on purpose: query the models, apply the grid, settle the doubtful cases, freeze the ranking. Changing the grid and rescoring costs no API call: raw answers are kept, which is what keeps comparisons over time honest.
The site itself calls no model. It reads frozen results, and refuses to build if a data file is invalid: better no site than a half-displayed ranking.
Acknowledged limits
One prompt per model, no fine-tuning, no specialised OCR upstream, images rather than PDFs with a text layer, and twenty-five documents rather than ten thousand. Published scores are a floor, not a ceiling.
Stating these limits before anyone throws them back at us is not modesty: it is what makes the rest credible.
Who publishes it
The hub is published by Flowera. It is meant for the executives, finance directors and operations managers of SMEs who have to choose an AI tool without being able to try it on their own documents.