We gave 10 open models
the same messy Monday morning
Eight files, seven of them tables. Two months of a SaaS company’s revenue, usage, churn, acquisition and support — and three decisions waiting on numbers that none of the tables contains.
One question: Can it reason across business data?
- customers.csv 36 rows customer_id segment, plan, ARR, renewal, status
- revenue.csv 71 rows customer_id invoices, two months, two sources
- usage.csv 67 rows usage_account_id requests, tokens, active users
- plans.csv 6 rows plan revenue, infra and support cost, included volume
- acquisition.csv 15 rows customer_id channel, new logo or expansion, attributed ARR
- funnel.csv 8 rows channel leads and paid trials, by channel and month
- support.csv 133 rows customer_id tickets, severity, status
usage_account_id joins on a different key — two accounts spell it another way
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models only All models
Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.
| # | Model | Monday Score | TTFA median | Show detail for | ||
|---|---|---|---|---|---|---|
| 1 1 | GLM 5.3 Flash glm5.3-flash | 98.0 | 766.2s | |||
Individual runs
Speed
Dimension breakdown
What it missed
What it got wrongNo penalties and no false actions in any run. | ||||||