// mondaybench · v1.7
Can a model handle
your Monday?
Five scenarios rebuild one real working morning. Every model gets the same inbox, the same contradictions and the same half-finished decisions, answers 3 times, and is scored against a frozen rubric.
// overall
Overall leaderboard
Monday Score is the equal-weight average across all five real-work scenarios.
Open models only All models
Leaderboard
9 open-weight models, ranked by Monday Score. 13 models ranked by Monday Score, including the 4 closed-weight models.
Quality band lower top
Models in one band share a shade. The suite does not order a band internally — the four at the top are within 0.54 points of each other against a minimum run sigma of 1.53 — so the bar length stays a true 0–100 reading and only the band changes colour.
| # | Model | Monday Score | Δ top | Worst Monday | Scenario spread |
|---|---|---|---|---|---|
| 01 | closed | 98.80 | +3.73 | 96.0 | 4.0 |
| 02 | closed | 97.07 | +2.00 −1.73 | 90.8 | 8.8 |
| 01 03 | 95.07 | −3.73 | 92.5 | 4.5 | |
| 02 04 | 94.93 | −0.14 −3.87 | 87.3 | 11.2 | |
| 05 | closed | 94.90 | −0.17 −3.90 | 89.3 | 10.7 |
| 03 06 | 94.53 | −0.54 −4.27 | 90.0 | 8.3 | |
| 04 07 | 93.73 | −1.34 −5.07 | 87.7 | 9.8 | |
| 05 08 | 91.70 | −3.37 −7.10 | 86.0 | 11.3 | |
| 09 | closed | 88.37 | −6.70 −10.43 | 78.5 | 21.5 |
| 06 10 | 86.67 | −8.40 −12.13 | 78.5 | 16.7 | |
| 07 11 | 86.50 | −8.57 −12.30 | 71.2 | 27.3 | |
| 08 12 | 68.77 | −26.30 −30.03 | 46.0 | 36.8 | |
| 09 13 | 63.00 | −32.07 −35.80 | 35.2 | 54.5 |
Expand scenario scores The five Mondays behind each average, shaded by score.
Cell shade 35.2 100.0
| Model | #001 | #002 | #003 | #004 | #005 | Scenario spread |
|---|---|---|---|---|---|---|
| closed | 100.0 | 100.0 | 100.0 | 98.0 | 96.0 | 4.0 |
| closed | 98.5 | 98.0 | 98.3 | 90.8 | 99.7 | 8.8 |
| 92.5 | 94.8 | 94.5 | 96.5 | 97.0 | 4.5 | |
| 98.0 | 87.3 | 92.8 | 98.0 | 98.5 | 11.2 | |
| closed | 97.5 | 91.0 | 96.7 | 89.3 | 100.0 | 10.7 |
| 95.0 | 93.0 | 96.3 | 90.0 | 98.3 | 8.3 | |
| 97.5 | 87.7 | 94.8 | 91.7 | 97.0 | 9.8 | |
| 90.8 | 89.3 | 95.0 | 86.0 | 97.3 | 11.3 | |
| closed | 97.0 | 81.0 | 85.3 | 78.5 | 100.0 | 21.5 |
| 88.2 | 81.2 | 90.3 | 78.5 | 95.2 | 16.7 | |
| 84.3 | 91.2 | 87.3 | 71.2 | 98.5 | 27.3 | |
| 82.8 | 70.2 | 65.5 | 46.0 | 79.3 | 36.8 | |
| 70.7 | 54.2 | 65.3 | 35.2 | 89.7 | 54.5 |
Technical details Run stability, unranked models, and what produced each number.
Run stability
How far apart the three official runs of one scenario land, averaged across the five. A model you cannot get the same answer from twice is a different product from one you can, so it is published here and never counted in the Monday Score.
| Model | Mean run range | Worst run range |
|---|---|---|
| closed | 0.3 | 1.5 |
| closed | 2.5 | 8.5 |
| closed | 3.5 | 9.5 |
| 4.3 | 10.5 | |
| closed | 5.4 | 23.5 |
| 6.0 | 19.0 | |
| 6.4 | 11.5 | |
| 6.8 | 13.5 | |
| 7.0 | 17.0 | |
| 9.1 | 16.0 | |
| 9.9 | 19.0 | |
| 14.7 | 31.0 | |
| 16.3 | 33.0 |
// cost & efficiency
Three models the benchmark cannot separate
The top three scores sit inside the measurement noise, so the ranking stops being the interesting question. What separates them is what they charge, how long they take and how much they write.
Quality vs cost
Opus records the highest measured quality, but at a very different price point from the leading open models. 97.07 at $246.91 per 1,000 tasks against 94.93 at $10.71.
Provider view mixes first-party list prices with public third-party provider prices.
From $0.44 to $456 per 1,000 tasks. More than 1037× difference in price across the field. A pricing range only — not a quality comparison.
Models with no published price stay at the end.
| Model | Price basis | Confidence | ||
|---|---|---|---|---|
| Claude Opus 5 | 97.07 | $246.91 | First-party | HIGH |
| Qwen 3.8 Flash On frontier | 95.07 | $12.61–$15.56 | First-party | MEDIUM |
| GLM 5.3 Flash On frontier | 94.93 | $10.71 · promo $5 | First-party | HIGH |
| Claude Fable 5.1 | 94.90 | $456.27 | First-party | HIGH |
| DeepSeek V4.1 Flash | 94.53 | $10.72–$22.47 | First-party | MEDIUM |
| GLM 5.3 | 93.73 | $99.20 | First-party identity contested | HIGH |
| DeepSeek V4 Flash | 91.70 | $12.94–$25.88 | First-party | HIGH |
| Gemini 3.8 Flash (High) | 88.37 | Unknown | — | — |
| MIMO v2.5 | 86.67 | $3.33 | OpenRouter | HIGH |
| GLM 5.2 | 86.50 | $39.50 | First-party | HIGH |
| Qwen 3.6 | 68.77 | $7.70 | OpenRouter ref | MEDIUM |
| Gemma 4 | 63.00 | $0.44 | OpenRouter ref | MEDIUM |
Cost at scale
| Model | 1k | 100k | 1M |
|---|---|---|---|
| Qwen 3.8 Flash | $13–$16 | $1.3k–$1.6k | $12.6k–$15.6k |
| GLM 5.3 Flash | $11 | $1.1k | $10.7k |
| Claude Fable 5.1 | $456 | $45.6k | $456k |
| Claude Opus 5 | $247 | $24.7k | $247k |
| Claude Fable 5.1 vs Qwen 3.8 Flash · 29.3×–36.2× | +$444 | +$44.4k | +$444k |
| Claude Fable 5.1 vs GLM 5.3 Flash · 42.6× | +$446 | +$44.6k | +$446k |
MondayBench-equivalent tasks. A linear extrapolation of the observed workload, not a universal per-request cost.
How each model gets to its result
Opus combines the highest measured score with the fastest first answer, but at a much higher cost than the leading open models.
Two observable dimensions. Neither is folded into Monday Score, and there is no efficiency score.
Speed
Fable reaches the same measured quality band roughly 5.08×–5.46× faster.
Concision
Average output tokens
Fable produces substantially shorter outputs on this workload.
2.45× fewer output tokens than GLM 5.3 Flash and 3.87× fewer than Qwen 3.8 Flash.
Observable generated output only. Provider observability is not homogeneous across models — detail only, not a comparison.
Not plotted: Gemini 3.8 Flash (High). The route reports no time-to-first-answer and no token count, and neither is estimated. Total generation latency for those runs is in the per-scenario tables.
// the suite
Your Monday, in five tests
Not five separate exams. One morning taken apart.
- #001 Understand Back to Work Can it figure out what actually matters?
- #002 Reconcile Moving Targets Can it tell what changed?
- #003 Plan Blockers Can it find the path through dependencies?
- #004 Analyse Numbers Don’t Lie Can it reason through messy business data?
- #005 Communicate Read the Room Can it tell four people the same truth?
// method
How MondayBench works
Every input, response, verdict and scoring rule is public.
How the suite was built
One morning, taken apart
Not five exams. Five situations out of the same Monday — an inbox, a schedule that moved, work blocked on other work, a pack of numbers, and a thing that has to be said carefully. Each one carries its own sources, its own contradictions and its own half-made decisions.
118 criteria, priced before the first run
Every scenario fixes its criteria and what each one is worth ahead of any generation, and the five sum to exactly 100 each. A later change opens a new rubric version rather than rewriting the one a published score was measured against; the drafts that lost stay on disk as the calibration record.
Frozen, then versioned
Each prompt is pinned by sha256, so a scenario that changed would stop matching its own hash. Runs, verdicts and boards can only be added to, never edited, and a release carries one digest over every artefact in it — clone the tag, recompute, and you get the same 64 characters or you have found something.
-
Same input
Identical scenario, source order and prompt for every model.
-
3 runs
Every official model answers three times.
-
Atomic rubric
Frozen criteria, each worth a stated number of points.
-
Recorded judge
Claude Opus 5 labels each criterion, with its evidence.
-
Deterministic score
Points add up. No weighting applied after the fact.
View GitHub Read the capability taxonomy
The judge is a record, not a call The five rules in full, what is guaranteed, and what is not.
- Every model receives the same scenario, source order and benchmark prompt.
- Every official model is run three times.
- Monday Score is the mean of those three official runs.
- Speed and cost are reported separately and never affect the quality score.
- Every input, raw response, judge verdict and scoring rule is public.
About the judge
The judge is Claude Opus 5 reading each response against the rubric in a recorded session, not an API call. That means a published result is reproducible as a record, since every label ships with its evidence and can be checked line by line, but a third party cannot regenerate the labels from scratch by running a command. The same effort that built the harness also produced the labels, which is a real bias risk rather than a hypothetical one. The audit trail is the mitigation, not a claim that none was needed.
// model fingerprints
Two models, where they differ
The same criteria as the leaderboard, regrouped by what they measure. Showing the 6 capabilities with the widest gap across the field, among those measured by at least 3 criteria: 174 of 500 rubric points.
Qwen 3.8 Flash
GLM 5.3 Flash
- temporal reasoning 3 crit · 14 pts gap 14
- supersession 6 crit · 27 pts gap 7
- causal discipline 8 crit · 23 pts gap 6
- uncertainty calibration 7 crit · 23 pts gap 5
- quantitative reasoning 16 crit · 66 pts gap 5
- business judgment 6 crit · 21 pts gap 0
Read a row, not a column: two models are comparable on one capability, two capabilities are not comparable on one model. The point mass is whatever the rubrics gave them, not a claim about importance.
// trust the result
Built to be audited
Auditable is not a claim about how the numbers came out. It is a set of things done before they existed, each of which someone outside this can check.
-
Nothing is discarded for scoring badly
Every official run is scored and published, the bad ones included. There is no best-of-N and no retry after a poor result, so the 3 runs behind a score are the 3 runs that happened.
-
The rubric is frozen before anything runs
Criteria and their points are fixed and versioned ahead of the first run. A later change opens a new rubric version rather than rewriting the one a published score was measured against.
-
Every verdict ships with its evidence
The judge labels one atomic criterion at a time, and each label carries the passage of the model’s own answer it was based on. A disagreement can be taken to the sentence it turns on.
-
Every model gets the identical input
Same scenario, same source order, same benchmark prompt. Nothing is tuned per model, and no model sees a running total while it answers.
-
The release is hashed and checkable from a clean clone
Each artefact carries a hash and the version carries one digest over all of them. Clone the tag, recompute, and you get the same digest or you have found something.
- Fresh-clone verified
- Byte-identical release
Technical release details The suite digest, and what it does and does not cover.
Verified from a clone of the tag rather than from a working copy: all five freeze manifests clean and every group digest matching. The digest below is the whole version in sixty-four characters.
Suite digest
060a2cdd79fcf6096039af5af4b11f4e96fabc683476c63579f01043d398ee45 This is a claim about the evidence trail, not about the judgement. The trail reproduces exactly; the generative labelling does not, and the methodology says so.
MondayBench v1.1 introduces scenario-specific evaluator contract fingerprints. Candidate generations and scoring semantics are unchanged from v1.0; existing recorded judgments were preserved, and every published score from v1.0 is identical here. GLM 5.3 is the first new model published under the v1.1 evaluator contract. v1.0 stays archived and byte-reproducible.
// why we built this
Benchmarks are useful
Production is the point
We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.