Skip to content
WorkBench.ai

ComingOn text

Meeting minutes

Do the minutes assign the right action to the right person?

The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.

What this test will measure

From a one-hour transcript, produce the decisions made and the actions with owner and deadline. A decision that was never made in the meeting is an invention, and counts as one.

What will be graded

  • Decisions
  • Actions and owners
  • Faithfulness, nothing added
  • Concision

What it is waiting for

No public reference says what the right answer is. A scoring rubric and human arbitration are needed before anything can be ranked.

Place on the roadmap
Wave 3
Target dataset
QMSum — 232 réunions transcrites et résumées par des humains ; la notation d'un résumé demande une grille
Target sample
40 meetings
What the model receives
Text

One business function at a time, one task at a time. Every wave ends with a publication.

No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.

How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.