Skip to content
WorkBench.ai

Analysis

Four measures, never a single score

Accuracy, cost, time, hallucinations: the hub never melts them into one number. A model can top the accuracy table and still be unusable because it makes things up.

A mark out of a hundred is convenient. It fits in a headline and it names a winner. It also hides what matters when the time comes to choose.

Take two models. The first reads an invoice slightly better than the second, but when a box is empty it sometimes fills it in. The second makes a few more mistakes and leaves empty what is empty. A weighted average would put the first one ahead. Your accountant would pick the second: an empty box gets noticed, whereas an invented VAT amount sails through review and into the books.

That is why the hub publishes four measures, side by side:

  • No review needed: the share of cases handled with nothing for a human to fix.
  • Accuracy: the points earned across all graded items, weighted by what a mistake costs.
  • Hallucinations: items invented where the source contains nothing.
  • Cost and time: per test case, at public prices on the day of the test; time is a median, so that one slow call does not skew everything.

They do not answer the same question. Accuracy tells you whether the model can do the job. The no-review rate tells you how much work is left for you. Hallucinations tell you whether you can trust it when nobody is looking. Cost and time tell you whether the whole thing still holds at ten thousand documents a month.

No formula can weigh those four answers for you. A practice handling three hundred invoices a month does not have the constraints of a platform receiving a hundred thousand: for one, price is a detail; for the other, it decides everything. How much each measure weighs depends on your case, and it is yours to set.

There is, all the same, one overall figure on the site: the Business Index. It aggregates only one of the four measures — accuracy. For each business function, we take the mean of the model's accuracy on that function's benchmarks; the index is the mean of those scores, every function weighing the same, whether it has two benchmarks or three.

So the index says nothing about cost, time or hallucinations, which stay in their own columns. Its job is to spot the models that can hold every post. To choose, open the leaderboard for the function that concerns you, and read all four columns.