We gave 10 open models
the same messy Monday morning
Slack messages. Emails. Meeting notes. A stale task list. Customer metrics.
One question: Can it figure out what actually matters?
metrics.csv Umbrella’s anomaly — in the metrics, mentioned by nobody
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models only All models
Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.
| # | Model | Monday Score | TTFA median | Show detail for |
|---|