Skip to content
WorkBench.ai

ComingOn documents

Decision memo

Forty pages of material, one page of memo: how much of it is still true?

The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.

What this test will measure

Write a one-page memo from a mixed file — studies, emails, spreadsheets. Graded by executives: are the key facts there, are the numbers right, does the recommendation follow from the file?

What will be graded

  • Key facts
  • Numerical accuracy
  • Options and risks
  • Reasoned recommendation

What it is waiting for

No public reference says what the right answer is. A scoring rubric and human arbitration are needed before anything can be ranked.

Place on the roadmap
Wave 3
Target sample
25 files
What the model receives
Documents (images or PDF)

One business function at a time, one task at a time. Every wave ends with a publication.

No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.

How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.