Skip to content
WorkBench.ai

Methodology

Two parts: the method common to the whole hub, then the protocol specific to each measured benchmark. Everything can be checked in the repository — the prompt sent, the scoring grid applied, the documents submitted and each model's raw answers.

The Business Index

The Business Index sums a model up in one figure. It is computed in two steps:

  • a business function's score is the mean of the model's accuracy on the benchmarks it took in that function, rounded to a tenth of a point;
  • the index is the mean of those per-function scores, rounded the same way.

So every business function weighs the same, however many benchmarks it has: a function with three tests does not count three times. A model scoring 90% and 70% on one function's two benchmarks, and 50% on another's only benchmark, has function scores of 80% and 50%. Its index is 65%, not 70%, the mean of the three tests.

The index covers only the business functions already measured. A function whose benchmarks have not been run does not enter the computation: it does not count as a zero, it does not count at all. How many functions the index covers is written next to it, because an index over one function is not worth an index over ten.

A model missing from a function the others took, however, has neither an index nor a rank: it appears as “Not ranked”, at the bottom of the tables. Its index could not be compared with any other.

A model that cannot read documents — neither images nor PDFs — does not take the benchmarks that submit one. Within each business function, it is scored on the text benchmarks alone; a function with none would leave it without an index. Its index therefore rests on fewer tests than the others': the number of benchmarks taken is shown alongside.

The index aggregates accuracy alone. A model's cost, time and hallucination rate are means over the benchmarks it took, shown separately and never melted into the index. An unknown price stays out of the cost mean.

Rank follows the index, from highest to lowest. On a benchmark, it follows accuracy.

The margin of error

A score measured on a few dozen documents is not known to a tenth of a point. The margin of error says so: it is the half-width of the 95% confidence interval on accuracy, in points. When a ranking publishes it, it is shown next to the score: “± 2.1”.

Two models separated by less than the larger of their two margins are not told apart. The table still puts them in order, because some order is needed; the benchmark page, for its part, states that the test does not separate them. With no published margin, the site declares no tie.

The Business Index has a margin of its own, combined from those of the benchmarks taken: the square root of the sum of their squares, divided by their number. It is published only if every benchmark the model took publishes one.

The margin describes the luck of the draw in the documents. It says nothing about the limits of the protocol, listed further down.

Costs and prices

The prices shown are the models' public prices, in dollars per million tokens, for input and output. They are synced from OpenRouter's public catalogue — the gateway the calls are actually billed through — as are context windows; last synced on 2 October 2026.

They stay in dollars: converting them would date the figure. A model missing from the catalogue is labelled “Undisclosed” rather than given a price copied from memory.

Cost per test is the mean of successful calls, at the price on the day of the test: the tokens each call consumed, multiplied by the catalogue price. It is recorded at run time, never reconstructed afterwards, and recomputed from published prices rather than read from the provider: anyone reading the repository can redo the multiplication.

A failed call enters no average, and is counted separately. A model that fails on more than a tenth of the set blocks publication of the ranking.

A failure may still have been billed: a model that answers and then exceeds the length limit consumes tokens without returning a usable result. Those calls are recorded with their cost, so that the total spent is accurate, but they do not weigh on a model's cost per test.

Failures do not penalise a model's score. In this project, the ones we observed came from the account quota, from credit reserved by concurrent calls, or from a host being down — from our side, that is. Their number stays in the ranking: a model failing for a reason of its own has to be visible.

The time shown is a median, not a mean: a single slow call would skew the mean.

How an answer is verified

Querying a model is the easy part. Everything else comes down to one question: how do we know its answer is right?

No AI grades an AI. The reference is written by humans, before the test and without knowing which models will take it. For the invoices, journalists entered every field by hand. Having one model judge another would measure their agreement, not their correctness.

The comparison is done by a hand-written program, field by field, according to the nature of the information:

  • an amount is reduced to a number, then compared to the cent;
  • a date is reduced to year-month-day from the dozen forms models use — “12/03/2020”, “March 12, 2020”, “2020-03-12” — applying the reading convention of the document's country, not the reader's;
  • an identifier or a name is compared after normalisation: accents, spaces, punctuation, common abbreviations;
  • a list of line items is matched line by line, order not counting.

This program never guesses. When it cannot decide — an unreadable value in the reference, a form it does not recognise — the answer goes to human review rather than being counted wrong. A reference we cannot read fails the preparation of the test: without that rule it would become a silent “nothing”, and a model that read the document correctly would be accused of inventing.

The same document, the same prompt, the same conditions for every model. A document one provider refuses — too many pages, too many pixels — is dropped from the test for everyone, not just for that provider: a ranking must compare models, not subsets of documents.

Ambiguous questions are excluded, not patched afterwards. If a field admits several defensible answers, scoring it measures whether the model guesses what we wanted, not whether it can read. The field is then removed from the rubric and the reason is written into the rubric itself; the models' answers stay published.

Finally, the pipeline's four stages are separate and immutable: call the models, score, arbitrate, publish. Raw answers are written once and for all, and re-scoring costs no call. That is what makes it possible to fix a rubric without ever touching an answer.

The raw answers of all 27 models queried are in the repository, one folder per model: anyone can redo the scoring and get the published figures back.

Nothing the site displays is fabricated. A task is measured and carries its figures, or it is on the roadmap and carries none.

As of today, no published ranking is a demonstration: every figure on the site comes from calls actually made.

Measured or coming

Every benchmark carries one of these two labels:

  • Runnable protocol: the task runs end to end and carries measured figures. Its definition, scoring grid, prompt, documents and annotations are in the repository; anyone can rerun the pipeline and get the same figures.
  • Coming: the protocol is written — the question asked, what the model receives, the subtasks, the target sample size — but the test set has still to be built. No figure is shown.

Measured as of today: 1 of 22 (Advertising invoice extraction).

A task that is coming states what it is waiting for, because the causes are not equivalent: a real, annotated public dataset still to be wired in; a task for which no public reference gives the right answer, and which therefore needs a scoring rubric and human arbitration; or documents that never leave the company, to be collected from partners with their consent.

Why publish a protocol before its results? Because it can no longer be adjusted afterwards to suit a ranking, and because it is easier to challenge before than after.

One benchmark's protocol

The protocol: Advertising invoice extraction

Everything above applies to the hub as a whole: the Business Index, the margin of error, where costs come from, how mature a task is. What follows applies to one benchmark only. Its limits, its scoring grid, its verdicts, its prompt and its latest run belong to it, and say nothing about the others.

Every benchmark will get its own section here as it is measured. A protocol does not carry over from one task to the next: an amount is checked to the cent, whereas sorting applications or summarising a contract calls for a different grid entirely — and for judgements no automatic comparator can make.

View the benchmark

Advertising invoice extraction

What this test does not measure

  • One prompt per model. A prompt tuned for a given model would improve its results; we measure what a standard integration gives.
  • No fine-tuning, no specialised OCR upstream. A dedicated processing chain would do better.
  • One hundred documents, not ten thousand. Enough to see a twenty-point gap, not to separate two models one point apart. Each score's margin of error says so.
  • American invoices, in English, from a single sector: television advertising buys. Nothing guarantees these results carry over to a French invoice.
  • Images, not PDFs with a text layer. A native PDF is easier for a model to read: published scores are a floor, not a ceiling.
  • Two scored fields. The other fields on these invoices admit several defensible answers — six competing identifiers, two distinct periods — and the question is no longer asked.
  • No line items. That is the real difficulty of the job, and the original annotation does not cover it: we would have to build our own reference.
  • A task easier than expected. Twenty-five models out of twenty-seven exceed 97%: this test separates them poorly, and should be read as an entry threshold rather than a ranking.

Advertising invoice extraction

The scoring grid

Fields are weighted by what a mistake costs in accounting, not by how technically hard they are. One critical field wrong drops the whole invoice into “needs review”.

Scoring grid for French invoice reading
FieldComparisonWeightCritical
Montant total facturéto the cent3yes
Numéro de TVAidentical, punctuation ignored1no

Field labels are the original French ones, as written in the task definition.

A field is worth its weight or zero, with no half points. Accuracy is the sum of points earned over the total possible. An invoice counts as “no review needed” when none of its critical fields is wrong, missing or invented; an invoice the model failed to process does not count.

The comparison is fussy where the business is fussy: amounts are compared to the cent. It is lenient where the business is lenient too: an identifier may be written with or without spaces, and the order of line items does not matter.

Dates follow the convention of the document's country, declared field by field in the rubric. On these American invoices, 03/04/2020 is read as 4 March; on a French invoice the same writing would mean 3 April. Reading an American date the French way had in fact produced, on a first attempt, some twenty errors wrongly blamed on the models.

Advertising invoice extraction

The three verdicts

A field absent from the document is part of the test. Answering “nothing” when there is nothing counts as a right answer; producing a value counts as a hallucination, tallied separately from the rest and never melted into an overall score.

  • Correct: the value on the document, or “nothing” when the document contains nothing.
  • Wrong or missing: a value other than the one on the document, or nothing when the information was there.
  • Hallucinated: a value where the document contains none.

Only a correct field earns points. A hallucination also feeds a measure of its own: the share of genuinely absent fields that the model filled in anyway. A model that writes “non trouvé” or “n/a” is abstaining correctly: its answer is reduced to “nothing” before it is scored.

The code does not have the last word. A human arbitration stage selects the answers to review — every hallucination, those judged correct but written differently from the reference, and a sample of the rest — and the human's verdict replaces the code's in the computation.

On the published run, that stage has not been played yet: the figures shown are the comparator's alone. 80 of the 2700 gradings are flagged “needs review”, and the ranking will say otherwise once they have been arbitrated. Saying so, rather than implying a review that did not happen, is part of the protocol.

Advertising invoice extraction

The prompt, in full

Sent as is to every model, along with the scanned pages of the invoice. The same for all twenty-seven, without a word of difference.

You are given the pages of a broadcast advertising invoice, as images.
Extract the requested information and answer with a JSON object only.

## The most important rule

If a piece of information **does not appear** on the document, answer `null` for
that field. Never guess, never infer, never fill in what looks usual. A value you
invent is treated as a serious error — worse than admitting you did not find it.

## Fields

- `contract_num` — the contract or order number identifying this buy, exactly as
  printed. String.
- `advertiser` — the value of the field labelled `Advertiser` or `Advertiser Name`:
  the client who bought the advertising. Not the TV station, not the media agency,
  and not the `Product` or `Brand` field, which often repeats the advertiser's name
  in a different form. String.
- `gross_amount` — the total gross amount invoiced, for the whole document.
  This is usually a grand total, and it is often on a later page than the first.
  Number.
- `flight_from` — the first day of the advertising period covered by this invoice.
  When the document names that period — a field labelled `Flight Dates`,
  `Order Flight` or `Flight` — use it, and prefer it over `Invoice Period`,
  `Bill Period` or `Bill Plan`, which cover the billing cycle and often start on a
  different day. **Many of these documents have no flight field at all**: they
  carry only a billing period, usually labelled `Period`. Use that one then, and
  do not answer `null` — the period is on the document, under another name.
  Format `YYYY-MM-DD`. String.
- `flight_to` — the last day of that flight. Format `YYYY-MM-DD`. String.
- `vat_number` — the VAT registration number of the issuing company, if the
  document carries one. String.

## Value formats

- Dates: `YYYY-MM-DD`. A date printed `02/03/20` on a US document means
  3 February 2020, so it becomes `2020-02-03`.
- Amounts: a plain decimal number, **no currency symbol and no thousands
  separator**. `$1,880.00` becomes `1880.00`.

## Output

A JSON object with exactly these six keys, and nothing around it.

Advertising invoice extraction

This test

Latest published run of invoice reading
Run2026-09-27_facture-fcc+2026-10-02_facture-fcc+2026-10-04_facture-fcc+2026-10-05_facture-fcc
Date5 October 2026
Documents100
Models27
Statusreal measurement
View the benchmark