Business Index
Can one model hold every desk? The index is the mean of its accuracy scores per business function, every function weighing the same.
Updated 5 October 2026 · Index computed over 1 of 10 business functions and 1 of 22 benchmarks.
- GPT-6 Luna100.0%
- DeepSeek V4.1 Flash100.0%
- Command A+100.0%
- Kimi K2.6100.0%
- Muse Spark 1.3100.0%
- Gemma 4 31B99.3%
- Qwen3.7 Flash99.3%
- Ministral 3 8B 251299.3%
- GLM 5.3 Flash99.3%
- Claude Sonnet 599.3%
- Grok 4.799.3%
- Nova Lite 1.084.3%
Each lab's best model, ranked by Business Index.
View full resultsKey takeaways
- GPT-6 Luna leads the index with 100.0%, ahead of DeepSeek V4.1 Flash (100.0%) and Command A+ (100.0%).
- Best open-weights model: DeepSeek V4.1 Flash, ranked 2 with 100.0%.
- Best model from a French lab: Ministral 3 8B 2512, ranked 9 with 99.3%.
The overall leaderboard
The index, then the score for each function. The best value in each column is highlighted.
| Rank | Model | Benchmarks | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Luna | 100.0%±0.0 | — | 100.0% | — | — | — | — | — | — | — | — | 1 / 22 |
| 2 | DeepSeek V4.1 Flash | 100.0%±0.0 | — | 100.0% | — | — | — | — | — | — | — | — | 1 / 22 |
| 3 | Command A+ | 100.0%±0.0 | — | 100.0% | — | — | — | — | — | — | — | — | 1 / 22 |
| 4 | Kimi K2.6 | 100.0%±0.0 | — | 100.0% | — | — | — | — | — | — | — | — | 1 / 22 |
| 5 | Muse Spark 1.3 | 100.0%±0.0 | — | 100.0% | — | — | — | — | — | — | — | — | 1 / 22 |
| 6 | Kimi K3 | 99.8%±0.7 | — | 99.8% | — | — | — | — | — | — | — | — | 1 / 22 |
| 7 | Gemma 4 31B | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 8 | Qwen3.7 Flash | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 9 | Ministral 3 8B 2512 | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 10 | GLM 5.3 Flash | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 11 | Gemini 3.5 Flash Lite | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 12 | Qwen3.8 27B | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 13 | Gemini 3.8 Flash | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 14 | Claude Sonnet 5 | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 15 | Qwen3.8 Max (0902) | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 16 | Grok 4.7 | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 17 | Claude Opus 5.5 | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 18 | Qwen3.8 Max Prime | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 19 | GPT-6 Astra | 99.3%±1.2 | — | 99.3% | — | — | — | — | — | — | — | — | 1 / 22 |
| 20 | GPT-6 Sol | 98.5%±1.7 | — | 98.5% | — | — | — | — | — | — | — | — | 1 / 22 |
How the index is computed
For each business function, we average the model's accuracy over the benchmarks of that function it was able to take. The index is the mean of those scores: a function with three benchmarks weighs no more than one with two.
The index covers only the functions already measured, and the header says which. A function on the roadmap does not count as a zero — it does not count at all. While a single function is measured, the index is worth no more than the benchmark it comes from, and should be read that way.
A model that cannot read documents is scored, within each function, on text benchmarks only. Its index stays comparable but rests on fewer tests: the “Benchmarks” column says so.
The index says nothing about cost, response time or hallucinations. Those measures stay separate, in each benchmark: a model that leads on accuracy can be unusable because it makes things up.