Skip to content
WorkBench.ai

Business Index

Can one model hold every desk? The index is the mean of its accuracy scores per business function, every function weighing the same.

Updated 5 October 2026 · Index computed over 1 of 10 business functions and 1 of 22 benchmarks.

Key takeaways

  • GPT-6 Luna leads the index with 100.0%, ahead of DeepSeek V4.1 Flash (100.0%) and Command A+ (100.0%).
  • Best open-weights model: DeepSeek V4.1 Flash, ranked 2 with 100.0%.
  • Best model from a French lab: Ministral 3 8B 2512, ranked 9 with 99.3%.

The overall leaderboard

The index, then the score for each function. The best value in each column is highlighted.

Business Index
RankModelBenchmarks
1GPT-6 Luna100.0%±0.0—100.0%————————1 / 22
2DeepSeek V4.1 Flash100.0%±0.0—100.0%————————1 / 22
3Command A+100.0%±0.0—100.0%————————1 / 22
4Kimi K2.6100.0%±0.0—100.0%————————1 / 22
5Muse Spark 1.3100.0%±0.0—100.0%————————1 / 22
6Kimi K399.8%±0.7—99.8%————————1 / 22
7Gemma 4 31B99.3%±1.2—99.3%————————1 / 22
8Qwen3.7 Flash99.3%±1.2—99.3%————————1 / 22
9Ministral 3 8B 251299.3%±1.2—99.3%————————1 / 22
10GLM 5.3 Flash99.3%±1.2—99.3%————————1 / 22
11Gemini 3.5 Flash Lite99.3%±1.2—99.3%————————1 / 22
12Qwen3.8 27B99.3%±1.2—99.3%————————1 / 22
13Gemini 3.8 Flash99.3%±1.2—99.3%————————1 / 22
14Claude Sonnet 599.3%±1.2—99.3%————————1 / 22
15Qwen3.8 Max (0902)99.3%±1.2—99.3%————————1 / 22
16Grok 4.799.3%±1.2—99.3%————————1 / 22
17Claude Opus 5.599.3%±1.2—99.3%————————1 / 22
18Qwen3.8 Max Prime99.3%±1.2—99.3%————————1 / 22
19GPT-6 Astra99.3%±1.2—99.3%————————1 / 22
20GPT-6 Sol98.5%±1.7—98.5%————————1 / 22

How the index is computed

For each business function, we average the model's accuracy over the benchmarks of that function it was able to take. The index is the mean of those scores: a function with three benchmarks weighs no more than one with two.

The index covers only the functions already measured, and the header says which. A function on the roadmap does not count as a zero — it does not count at all. While a single function is measured, the index is worth no more than the benchmark it comes from, and should be read that way.

A model that cannot read documents is scored, within each function, on text benchmarks only. Its index stays comparable but rests on fewer tests: the “Benchmarks” column says so.

The index says nothing about cost, response time or hallucinations. Those measures stay separate, in each benchmark: a model that leads on accuracy can be unusable because it makes things up.

Methodology