// mondaybench · v1.5
Can a model handle
your Monday?
Five scenarios rebuild one real working morning. Every model gets the same inbox, the same contradictions and the same half-finished decisions, answers 3 times, and is scored against a frozen rubric.
// overall
Overall leaderboard
Monday Score is the equal-weight average across all five real-work scenarios.
Open models only All models
Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.
Leaderboard
9 open-weight models, ranked by Monday Score. 13 models ranked by Monday Score, including the 4 closed-weight models.
| # | Model | Monday Score | Δ top | Worst Monday | Scenario spread |
|---|---|---|---|---|---|
| 01 | closed | 98.80 | +3.73 | 96.0 | 4.0 |
| 02 | closed | 97.07 | +2.00 −1.73 | 90.8 | 8.8 |
| 01 03 | 95.07 | −3.73 | 92.5 | 4.5 | |
| 02 04 | 94.93 | −0.14 −3.87 | 87.3 | 11.2 | |
| 05 | closed | 94.90 | −0.17 −3.90 | 89.3 | 10.7 |
| 03 06 | 94.53 | −0.54 −4.27 | 90.0 | 8.3 | |
| 04 07 | 93.73 | −1.34 −5.07 | 87.7 | 9.8 | |
| 05 08 | 91.70 | −3.37 −7.10 | 86.0 | 11.3 | |
| 09 | closed | 88.37 | −6.70 −10.43 | 78.5 | 21.5 |
| 06 10 | 86.67 | −8.40 −12.13 | 78.5 | 16.7 | |
| 07 11 | 86.50 | −8.57 −12.30 | 71.2 | 27.3 | |
| 08 12 | 68.77 | −26.30 −30.03 | 46.0 | 36.8 | |
| 09 13 | 63.00 | −32.07 −35.80 | 35.2 | 54.5 |
Expand scenario scores The five Mondays behind each average, shaded by score.
Cell shade 35.2 100.0
| Model | #001 | #002 | #003 | #004 | #005 | Scenario spread |
|---|---|---|---|---|---|---|
| closed | 100.0 | 100.0 | 100.0 | 98.0 | 96.0 | 4.0 |
| closed | 98.5 | 98.0 | 98.3 | 90.8 | 99.7 | 8.8 |
| 92.5 | 94.8 | 94.5 | 96.5 | 97.0 | 4.5 | |
| 98.0 | 87.3 | 92.8 | 98.0 | 98.5 | 11.2 | |
| closed | 97.5 | 91.0 | 96.7 | 89.3 | 100.0 | 10.7 |
| 95.0 | 93.0 | 96.3 | 90.0 | 98.3 | 8.3 | |
| 97.5 | 87.7 | 94.8 | 91.7 | 97.0 | 9.8 | |
| 90.8 | 89.3 | 95.0 | 86.0 | 97.3 | 11.3 | |
| closed | 97.0 | 81.0 | 85.3 | 78.5 | 100.0 | 21.5 |
| 88.2 | 81.2 | 90.3 | 78.5 | 95.2 | 16.7 | |
| 84.3 | 91.2 | 87.3 | 71.2 | 98.5 | 27.3 | |
| 82.8 | 70.2 | 65.5 | 46.0 | 79.3 | 36.8 | |
| 70.7 | 54.2 | 65.3 | 35.2 | 89.7 | 54.5 |
Technical details Run stability, unranked models, and what produced each number.
Run stability
How far apart the three official runs of one scenario land, averaged across the five. A model you cannot get the same answer from twice is a different product from one you can, so it is published here and never counted in the Monday Score.
| Model | Mean run range | Worst run range |
|---|---|---|
| closed | 0.3 | 1.5 |
| closed | 2.5 | 8.5 |
| closed | 3.5 | 9.5 |
| 4.3 | 10.5 | |
| closed | 5.4 | 23.5 |
| 6.0 | 19.0 | |
| 6.4 | 11.5 | |
| 6.8 | 13.5 | |
| 7.0 | 17.0 | |
| 9.1 | 16.0 | |
| 9.9 | 19.0 | |
| 14.7 | 31.0 | |
| 16.3 | 33.0 |
// what it costs
Three models the benchmark cannot separate
The top three scores sit inside the measurement noise, so the ranking stops being the interesting question. What separates them is what they charge, how long they take and how much they write.
Top performance band
-
Highest in band
Qwen 3.8 Flash
95.07 / 100
$12.61–$15.56 / 1k · 258.9s
-
Cheapest top-band
GLM 5.3 Flash
94.93 / 100
$10.71 / 1k
-
Most concise top tier
Claude Fable 5.1
94.90 / 100
50.9s · ~8.2k output
Fable reaches the same measured quality band roughly 5.08×–5.46× faster while generating 2.45×–3.87× fewer output tokens.
The trade-off: 29.3×–42.6× the cost.
Quality vs cost
Opus records the highest measured quality, but at a very different price point from the leading open models. 97.07 at $246.91 per 1,000 tasks against 94.93 at $10.71.
GLM 5.3 appears twice: $99.20 per 1,000 tasks at the vendor’s list price, $5.67 on the metered route the benchmark actually ran. Same model, same score, 17.5× apart — model price and route price can tell very different stories.
Provider view mixes first-party list prices with public third-party provider prices.
From $0.44 to $456 per 1,000 tasks. More than 1037× difference in price across the field. A pricing range only — not a quality comparison.
| Model | Monday Score | Cost / 1k | Price basis | Confidence |
|---|---|---|---|---|
| Claude Opus 5 | 97.07 | $246.91 | First-party | HIGH |
| Qwen 3.8 Flash On frontier | 95.07 | $12.61–$15.56 | First-party | MEDIUM |
| GLM 5.3 Flash On frontier | 94.93 | $10.71 · promo $5 | First-party | HIGH |
| Claude Fable 5.1 | 94.90 | $456.27 | First-party | HIGH |
| DeepSeek V4.1 Flash | 94.53 | $10.72–$22.47 | First-party | MEDIUM |
| GLM 5.3 | 93.73 | $99.20 $5.67 · route | First-party identity contested Helmcode route | HIGH |
| DeepSeek V4 Flash | 91.70 | $12.94–$25.88 | First-party | HIGH |
| Gemini 3.8 Flash (High) | 88.37 | Unknown | — | — |
| MIMO v2.5 | 86.67 | $3.33 | OpenRouter | HIGH |
| GLM 5.2 | 86.50 | $39.50 | First-party | HIGH |
| Qwen 3.6 | 68.77 | $7.70 | OpenRouter ref | MEDIUM |
| Gemma 4 | 63.00 | $0.44 | OpenRouter ref | MEDIUM |
Cost at scale
| Model | 1k | 100k | 1M |
|---|---|---|---|
| Qwen 3.8 Flash | $13–$16 | $1.3k–$1.6k | $12.6k–$15.6k |
| GLM 5.3 Flash | $11 | $1.1k | $10.7k |
| Claude Fable 5.1 | $456 | $45.6k | $456k |
| Claude Opus 5 | $247 | $24.7k | $247k |
| Claude Fable 5.1 vs Qwen 3.8 Flash · 29.3×–36.2× | +$444 | +$44.4k | +$444k |
| Claude Fable 5.1 vs GLM 5.3 Flash · 42.6× | +$446 | +$44.6k | +$446k |
MondayBench-equivalent tasks. A linear extrapolation of the observed workload, not a universal per-request cost.
How each model gets to its result
Opus combines the highest measured score with the fastest first answer, but at a much higher cost than the leading open models.
Two observable dimensions. Neither is folded into Monday Score, and there is no efficiency score.
Speed
Fable reaches the same measured quality band roughly 5.08×–5.46× faster.
Concision
Average output tokens
Fable produces substantially shorter outputs on this workload.
2.45× fewer output tokens than GLM 5.3 Flash and 3.87× fewer than Qwen 3.8 Flash.
Observable generated output only. Provider observability is not homogeneous across models — detail only, not a comparison.
Not plotted: Gemini 3.8 Flash (High). The route reports no time-to-first-answer and no token count, and neither is estimated. Total generation latency for those runs is in the per-scenario tables.
// what the table hides
Three things the ranking does not say
-
4 / 5
scenarios won outright by the model ranked first
The winner only wins one scenario
GPT-6 Astra ranks #1 because it never has a bad Monday.
Worst scenario: 96.0
-
98.8 ≠ 95.1
two scores, two different machines
Almost the same score. Very different models
GPT-6 has the higher floor. Qwen reaches higher peaks.
Range across scenarios: 4.0 vs 4.5
-
6 / 10
models score their worst on #004
Business data is where models break
Numbers Don’t Lie is the worst-scoring scenario for 6 of 10 models.
Widest field on the board
// the suite
Your Monday, in five tests
Not five separate exams. One morning taken apart.
- #001 Understand Back to Work Can it figure out what actually matters?
- #002 Reconcile Moving Targets Can it tell what changed?
- #003 Plan Blockers Can it find the path through dependencies?
- #004 Analyse Numbers Don’t Lie Can it reason through messy business data?
- #005 Communicate Read the Room Can it tell four people the same truth?
// the finding
Models can keep a secret
They don’t always know who needs it
0 / 30
explicit external leaks
11 / 30
failed to route the security item to engineering
The same sensitive fact had to stay away from the customer and the public status page, while still reaching Engineering.
// model fingerprints
Two models, where they differ
The same criteria as the leaderboard, regrouped by what they measure. Showing the 6 capabilities with the widest gap across the field, among those measured by at least 3 criteria: 174 of 500 rubric points.
Qwen 3.8 Flash
GLM 5.3 Flash
- temporal reasoning 3 crit · 14 pts gap 14
- supersession 6 crit · 27 pts gap 7
- causal discipline 8 crit · 23 pts gap 6
- uncertainty calibration 7 crit · 23 pts gap 5
- quantitative reasoning 16 crit · 66 pts gap 5
- business judgment 6 crit · 21 pts gap 0
Read a row, not a column: two models are comparable on one capability, two capabilities are not comparable on one model. The point mass is whatever the rubrics gave them, not a claim about importance.
// method
How MondayBench works
Every input, response, verdict and scoring rule is public.
-
Same input
Identical scenario, source order and prompt for every model.
-
3 runs
Every official model answers three times.
-
Atomic rubric
Frozen criteria, each worth a stated number of points.
-
Recorded judge
Claude Opus 5 labels each criterion, with its evidence.
-
Deterministic score
Points add up. No weighting applied after the fact.
View GitHub Read the capability taxonomy
The judge is a record, not a call The five rules in full, what is guaranteed, and what is not.
- Every model receives the same scenario, source order and benchmark prompt.
- Every official model is run three times.
- Monday Score is the mean of those three official runs.
- Speed and cost are reported separately and never affect the quality score.
- Every input, raw response, judge verdict and scoring rule is public.
About the judge
The judge is Claude Opus 5 reading each response against the rubric in a recorded session, not an API call. That means a published result is reproducible as a record, since every label ships with its evidence and can be checked line by line, but a third party cannot regenerate the labels from scratch by running a command. The same effort that built the harness also produced the labels, which is a real bias risk rather than a hypothetical one. The audit trail is the mitigation, not a claim that none was needed.
// trust the result
Built to be audited
Auditable is not a claim about how the numbers came out. It is a set of things done before they existed, each of which someone outside this can check.
-
Nothing is discarded for scoring badly
Every official run is scored and published, the bad ones included. There is no best-of-N and no retry after a poor result, so the 3 runs behind a score are the 3 runs that happened.
-
The rubric is frozen before anything runs
Criteria and their points are fixed and versioned ahead of the first run. A later change opens a new rubric version rather than rewriting the one a published score was measured against.
-
Every verdict ships with its evidence
The judge labels one atomic criterion at a time, and each label carries the passage of the model’s own answer it was based on. A disagreement can be taken to the sentence it turns on.
-
Every model gets the identical input
Same scenario, same source order, same benchmark prompt. Nothing is tuned per model, and no model sees a running total while it answers.
-
The release is hashed and checkable from a clean clone
Each artefact carries a hash and the version carries one digest over all of them. Clone the tag, recompute, and you get the same digest or you have found something.
- Fresh-clone verified
- Byte-identical release
Technical release details The suite digest, and what it does and does not cover.
Verified from a clone of the tag rather than from a working copy: all five freeze manifests clean and every group digest matching. The digest below is the whole version in sixty-four characters.
Suite digest
40c55b6c4b7cf5c597ee81edc5b26d416c1db92e949bb817766a2aeb5c9e3ba3 This is a claim about the evidence trail, not about the judgement. The trail reproduces exactly; the generative labelling does not, and the methodology says so.
MondayBench v1.1 introduces scenario-specific evaluator contract fingerprints. Candidate generations and scoring semantics are unchanged from v1.0; existing recorded judgments were preserved, and every published score from v1.0 is identical here. GLM 5.3 is the first new model published under the v1.1 evaluator contract. v1.0 stays archived and byte-reproducible.
// why we built this
Benchmarks are useful
Production is the point
We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.