Skip to content
WorkBench.ai

Which AI model for which business function?

Public leaderboards measure exams. We test the leading models on a company's real work — reading an invoice, summarising a contract, answering a customer — and publish the protocol, the documents and the raw answers.

27 models · 12 labs · 22 benchmarks across 10 business functions · 1 measured so far

Latest reports

Newly evaluated models, new benchmarks, notes on method.

All news

Benchmark5 Oct 2026

First measurement: 27 models read 100 real invoices

The winner costs $0.0009 per invoice and misses none out of a hundred. The two most expensive models in the panel miss one — and three models invent a VAT number that does not exist.

  • Five models out of twenty-seven read the total amount without a single mistake across one hundred invoices.
  • The winner, GPT-6 Luna, costs $0.0009 per invoice — eighty times less than GPT-6 Astra, from the same provider, which misses one.
  • Three models invent. On a VAT number these American invoices do not carry, Kimi K3, Llama 4 Maverick and Mistral Medium each fabricated a value. The other twenty-four abstain correctly.
Read the article

The leaderboard, function by function

The best model for finance is not necessarily the right one for your customer service. Pick a function, then a benchmark.

Advertising invoice extraction

How many invoices does a model read with nothing for a human to fix?

Each square is a model, in its lab's colour and monogram. Click to open its page.

The 24 most recent models. The green line links those that no other model beats on both price and accuracy.

21 further tasks, across 9 business functions, have a written protocol and are waiting for their test set. See what is on the roadmap

Performance over time

Each release that beats its group's record adds a step. Between two releases, the best available model does not change.

Business Index5 Oct 2026

Business Index of the best model released at each date. One point per release that improves the group's record.