// mondaybench #002 beta

We gave 10 open models
the same messy Monday morning

Nine sources. A release that failed on Friday, two ways to recover it, and four people whose calendars barely overlap.

One question: Can it tell what is still true?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for