// mondaybench #001 beta

We gave 10 open models
the same messy Monday morning

Slack messages. Emails. Meeting notes. A stale task list. Customer metrics.

One question: Can it figure out what actually matters?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
you log off wed 17:30 you’re back mon 09:07 thursday friday Release moves to Wednesday Launch email leaves your plate Contoso needs confirmation by 11:00
48 Slack messages · 7 emails · 2 meetings three of them change what is already true

metrics.csv Umbrella’s anomaly — in the metrics, mentioned by nobody

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for