ComingOn text
Business question to SQL
“Revenue by region this quarter”: is the query right?
The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.
What this test will measure
Turn a business question into a query over a realistic company schema, badly named tables and three date formats included. The query is executed: only the result counts.
What will be graded
- Joins
- Aggregations
- Dates and periods
- Ambiguous requests
What it is waiting for
A real, annotated public dataset exists. It remains to be wired into the pipeline.
- Place on the roadmap
- Wave 2
- Target dataset
- BIRD — 12 751 questions et requêtes de référence sur 95 bases réelles, CC BY-SA 4.0
- Target sample
- 70 questions
- What the model receives
- Text
One business function at a time, one task at a time. Every wave ends with a publication.
No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.
How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.