// mondaybench #004 beta

We gave 10 open models
the same messy Monday morning

Eight files, seven of them tables. Two months of a SaaS company’s revenue, usage, churn, acquisition and support — and three decisions waiting on numbers that none of the tables contains.

One question: Can it reason across business data?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
seven tables
  • customers.csv 36 rows customer_id segment, plan, ARR, renewal, status
  • revenue.csv 71 rows customer_id invoices, two months, two sources
  • usage.csv 67 rows usage_account_id requests, tokens, active users
  • plans.csv 6 rows plan revenue, infra and support cost, included volume
  • acquisition.csv 15 rows customer_id channel, new logo or expansion, attributed ARR
  • funnel.csv 8 rows channel leads and paid trials, by channel and month
  • support.csv 133 rows customer_id tickets, severity, status

usage_account_id joins on a different key — two accounts spell it another way

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 1 GLM 5.3 Flash glm5.3-flash 98.0 1.5 766.2s 827.8s