// mondaybench #003 beta

We gave 10 open models
the same messy Monday morning

Six sources. Six dated commitments across one week, three of them waiting on the same three-hour migration, and one engineer claimed by two of them at once.

One question: Can it plan through dependencies?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable

// six dated commitments · one week · 3 waits on the billing migration

Mon Tue Wed Thu Fri Contoso Q3 compliance export Julia Tue 16:00 Enterprise SSO live for Atlas Marcus Wed 17:00 Billing platform Q3 milestone Elena Thu 15:00 Public pricing page Sarah Thu 17:00 Board pack, revenue section Sarah Fri 10:00 Self-serve onboarding launch Julia Fri 12:00
  • now · Monday 09:10
  • waits on the billing migration
  • owes it nothing
01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for