We gave 10 open models
the same messy Monday morning
Nine sources. A release that failed on Friday, two ways to recover it, and four people whose calendars barely overlap.
One question: Can it tell what is still true?
// nine sources · 7.3 KB
- slack.md 69 lines Six channels, two days
- emails.md 18 lines Six emails
- meetings.md 14 lines Two meetings
- tasks.md 11 lines Eight tasks, five of them stale
- release-options.csv 2 rows The two recovery paths
- calendars.csv 5 rows Who is free, and when · decides the release
- accounts.csv 6 rows Usage, admin logins, open tickets
- contracts.csv 6 rows Renewal windows and commercial context
- support.csv 6 rows Six open tickets
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models only All models
Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.
| # | Model | Monday Score | TTFA median | Show detail for |
|---|