Skip to content
WorkBench.ai

Announcement

The hub opens up to ten business functions

From invoice reading alone to 22 benchmarks, from finance to management. One is measured today; the other twenty-one have a written protocol and a test set still to be built.

The hub began with a single question: how many French invoices can a model read without a human having to fix anything? It was the right place to start, because an invoice can be graded without argument: the amount is right or it is not. But a company is more than its accounts payable.

The hub now covers ten business functions: finance, accounting, human resources, legal, sales, marketing, customer service, procurement and logistics, IT, management. Twenty-two benchmarks in all, two or three per function.

Each one starts from a question an executive actually asks, not from some abstract capability of the model:

  • Does the model see the auto-renewal buried on page twelve?
  • Does the model promise a refund your terms do not allow?
  • Is the cheapest on paper still cheapest once shipping and penalties are in?

What is measured today: invoice reading, on real television advertising invoices filed with the US telecom regulator and annotated by hand. The prompt, the documents, the annotations, the scoring rubric and the scoring code are in the repository, and the pipeline runs the task end to end.

What is not measured yet: the other twenty-one benchmarks. Their protocol is written — the question, what the model receives, the subtasks, the target sample size — but their test set remains to be built. Their page is marked “Coming” and states what it is waiting for: a public dataset to wire in, a scoring rubric to write, or documents only companies hold.

Why show them anyway? Because a protocol published before the results cannot be quietly adjusted afterwards to suit a ranking. And because it is easier to challenge before than after: if we are asking a business function the wrong question, we would rather hear it now.

The hub displays no fabricated score. A task is measured and carries its figures, or it is on the roadmap and carries none. There is no third state.

To read the whole, there is the Business Index: the mean of the per-function scores, every function weighing the same. It replaces no leaderboard. The best model for finance is not necessarily the right one for your customer service, and that is exactly what the hub sets out to show.