// mondaybench · v1.7

Can a model handle
your Monday?

Five scenarios rebuild one real working morning. Every model gets the same inbox, the same contradictions and the same half-finished decisions, answers 3 times, and is scored against a frozen rubric.

// overall

Overall leaderboard

Monday Score is the equal-weight average across all five real-work scenarios.

Open models only All models

Leaderboard

9 open-weight models, ranked by Monday Score. 13 models ranked by Monday Score, including the 4 closed-weight models.

Quality band lower top

Models in one band share a shade. The suite does not order a band internally — the four at the top are within 0.54 points of each other against a minimum run sigma of 1.53 — so the bar length stays a true 0–100 reading and only the band changes colour.

# Model Monday Score Δ top Worst Monday Scenario spread
01 closed
98.80
+3.73 96.0
4.0
02 closed
97.07
+2.00 −1.73 90.8
8.8
01 03
95.07
−3.73 92.5
4.5
02 04
94.93
−0.14 −3.87 87.3
11.2
05 closed
94.90
−0.17 −3.90 89.3
10.7
03 06
94.53
−0.54 −4.27 90.0
8.3
04 07
93.73
−1.34 −5.07 87.7
9.8
05 08
91.70
−3.37 −7.10 86.0
11.3
09 closed
88.37
−6.70 −10.43 78.5
21.5
06 10
86.67
−8.40 −12.13 78.5
16.7
07 11
86.50
−8.57 −12.30 71.2
27.3
08 12
68.77
−26.30 −30.03 46.0
36.8
09 13
63.00
−32.07 −35.80 35.2
54.5
Expand scenario scores The five Mondays behind each average, shaded by score.

Cell shade 35.2 100.0

Model #001 #002 #003 #004 #005 Scenario spread
closed 100.0 100.0 100.0 98.0 96.0 4.0
closed 98.5 98.0 98.3 90.8 99.7 8.8
92.5 94.8 94.5 96.5 97.0 4.5
98.0 87.3 92.8 98.0 98.5 11.2
closed 97.5 91.0 96.7 89.3 100.0 10.7
95.0 93.0 96.3 90.0 98.3 8.3
97.5 87.7 94.8 91.7 97.0 9.8
90.8 89.3 95.0 86.0 97.3 11.3
closed 97.0 81.0 85.3 78.5 100.0 21.5
88.2 81.2 90.3 78.5 95.2 16.7
84.3 91.2 87.3 71.2 98.5 27.3
82.8 70.2 65.5 46.0 79.3 36.8
70.7 54.2 65.3 35.2 89.7 54.5
Technical details Run stability, unranked models, and what produced each number.

Run stability

How far apart the three official runs of one scenario land, averaged across the five. A model you cannot get the same answer from twice is a different product from one you can, so it is published here and never counted in the Monday Score.

Model Mean run range Worst run range
closed 0.3 1.5
closed 2.5 8.5
closed 3.5 9.5
4.3 10.5
closed 5.4 23.5
6.0 19.0
6.4 11.5
6.8 13.5
7.0 17.0
9.1 16.0
9.9 19.0
14.7 31.0
16.3 33.0

// cost & efficiency

Three models the benchmark cannot separate

The top three scores sit inside the measurement noise, so the ranking stops being the interesting question. What separates them is what they charge, how long they take and how much they write.

Result views

Quality vs cost

Opus records the highest measured quality, but at a very different price point from the leading open models. 97.07 at $246.91 per 1,000 tasks against 94.93 at $10.71.

Provider view mixes first-party list prices with public third-party provider prices.

60 70 80 90 100 $0.10 $1 $10 $100 $1000 Cost per 1,000 MondayBench-equivalent tasks · Log scale Claude Opus 5 — 97.07 · $246.91/1k · First-party · HIGH Claude Opus 5 Qwen 3.8 Flash — 95.07 · $12.61–$15.56/1k · First-party · MEDIUM Qwen 3.8 Flash GLM 5.3 Flash — 94.93 · $10.71/1k · First-party · HIGH GLM 5.3 Flash Claude Fable 5.1 — 94.90 · $456.27/1k · First-party · HIGH Claude Fable 5.1 DeepSeek V4.1 Flash — 94.53 · $10.72–$22.47/1k · First-party · MEDIUM DeepSeek V4.1 Flash GLM 5.3 — 93.73 · $99.20/1k · First-party · HIGH GLM 5.3 DeepSeek V4 Flash — 91.70 · $12.94–$25.88/1k · First-party · HIGH DeepSeek V4 Flash MIMO v2.5 — 86.67 · $3.33/1k · OpenRouter · HIGH MIMO v2.5 GLM 5.2 — 86.50 · $39.50/1k · First-party · HIGH GLM 5.2 Qwen 3.6 — 68.77 · $7.70/1k · OpenRouter · MEDIUM Qwen 3.6 Gemma 4 — 63.00 · $0.44/1k · OpenRouter · MEDIUM Gemma 4
First-party only OpenRouter Helmcode route Best quality for the price · vendor list Best quality for the price · route

From $0.44 to $456 per 1,000 tasks. More than 1037× difference in price across the field. A pricing range only — not a quality comparison.

Models with no published price stay at the end.

Quality vs cost
Model Price basis Confidence
Claude Opus 5 97.07 $246.91 First-party HIGH
Qwen 3.8 Flash On frontier 95.07 $12.61–$15.56 First-party MEDIUM
GLM 5.3 Flash On frontier 94.93 $10.71 · promo $5 First-party HIGH
Claude Fable 5.1 94.90 $456.27 First-party HIGH
DeepSeek V4.1 Flash 94.53 $10.72–$22.47 First-party MEDIUM
GLM 5.3 93.73 $99.20 First-party identity contested HIGH
DeepSeek V4 Flash 91.70 $12.94–$25.88 First-party HIGH
Gemini 3.8 Flash (High) 88.37 Unknown — —
MIMO v2.5 86.67 $3.33 OpenRouter HIGH
GLM 5.2 86.50 $39.50 First-party HIGH
Qwen 3.6 68.77 $7.70 OpenRouter ref MEDIUM
Gemma 4 63.00 $0.44 OpenRouter ref MEDIUM
Cost at scale
Model 1k 100k 1M
Qwen 3.8 Flash $13–$16$1.3k–$1.6k$12.6k–$15.6k
GLM 5.3 Flash $11$1.1k$10.7k
Claude Fable 5.1 $456$45.6k$456k
Claude Opus 5 $247$24.7k$247k
Claude Fable 5.1 vs Qwen 3.8 Flash · 29.3×–36.2× +$444+$44.4k+$444k
Claude Fable 5.1 vs GLM 5.3 Flash · 42.6× +$446+$44.6k+$446k

MondayBench-equivalent tasks. A linear extrapolation of the observed workload, not a universal per-request cost.

How each model gets to its result

Opus combines the highest measured score with the fastest first answer, but at a much higher cost than the leading open models.

Two observable dimensions. Neither is folded into Monday Score, and there is no efficiency score.

Speed

60 70 80 90 100 0s60s120s180s240s300s Median time to first answer Qwen 3.8 Flash — 95.07 · 258.9s · ~31.7k output Qwen 3.8 Flash GLM 5.3 Flash — 94.93 · 278.2s · ~20.1k output GLM 5.3 Flash Claude Fable 5.1 — 94.90 · 50.9s · ~8.2k output Claude Fable 5.1 DeepSeek V4.1 Flash — 94.53 · 69.2s · ~17.5k output DeepSeek V4.1 Flash GLM 5.3 — 93.73 · 253.3s · ~21.3k output GLM 5.3 DeepSeek V4 Flash — 91.70 · 121.4s · ~18.2k output DeepSeek V4 Flash MIMO v2.5 — 86.67 · 170.4s · ~13.1k output MIMO v2.5 GLM 5.2 — 86.50 · 44.1s · ~8.1k output GLM 5.2 Qwen 3.6 — 68.77 · 1.0s · ~3.3k output Qwen 3.6 Gemma 4 — 63.00 · 0.6s · ~0.9k output Gemma 4 fastest measured Claude Opus 5 — 97.07 · 41.3s · ~9.2k output Claude Opus 5

Fable reaches the same measured quality band roughly 5.08×–5.46× faster.

Concision

  • Gemma 4 ~0.9k Provider-reported reasoning tokens: ~0.0k
  • Qwen 3.6 ~3.3k Provider-reported reasoning tokens: ~1.2k
  • GLM 5.2 ~8.1k Provider-reported reasoning tokens: ~5.9k
  • Claude Fable 5.1 ~8.2k Provider-reported reasoning tokens: ~4.4k
  • Claude Opus 5 ~9.2k Provider-reported reasoning tokens: ~4.8k
  • MIMO v2.5 ~13.1k Provider-reported reasoning tokens: ~0.0k
  • DeepSeek V4.1 Flash ~17.5k Provider-reported reasoning tokens: ~15.6k
  • DeepSeek V4 Flash ~18.2k Provider-reported reasoning tokens: ~16.5k
  • GLM 5.3 Flash ~20.1k Provider-reported reasoning tokens: ~18.0k
  • GLM 5.3 ~21.3k Provider-reported reasoning tokens: ~19.3k
  • Qwen 3.8 Flash ~31.7k Provider-reported reasoning tokens: ~27.4k

Average output tokens

Fable produces substantially shorter outputs on this workload.

2.45× fewer output tokens than GLM 5.3 Flash and 3.87× fewer than Qwen 3.8 Flash.

Observable generated output only. Provider observability is not homogeneous across models — detail only, not a comparison.

Not plotted: Gemini 3.8 Flash (High). The route reports no time-to-first-answer and no token count, and neither is estimated. Total generation latency for those runs is in the per-scenario tables.

// method

How MondayBench works

Every input, response, verdict and scoring rule is public.

How the suite was built

One morning, taken apart

Not five exams. Five situations out of the same Monday — an inbox, a schedule that moved, work blocked on other work, a pack of numbers, and a thing that has to be said carefully. Each one carries its own sources, its own contradictions and its own half-made decisions.

118 criteria, priced before the first run

Every scenario fixes its criteria and what each one is worth ahead of any generation, and the five sum to exactly 100 each. A later change opens a new rubric version rather than rewriting the one a published score was measured against; the drafts that lost stay on disk as the calibration record.

Frozen, then versioned

Each prompt is pinned by sha256, so a scenario that changed would stop matching its own hash. Runs, verdicts and boards can only be added to, never edited, and a release carries one digest over every artefact in it — clone the tag, recompute, and you get the same 64 characters or you have found something.

  1. Same input

    Identical scenario, source order and prompt for every model.

  2. 3 runs

    Every official model answers three times.

  3. Atomic rubric

    Frozen criteria, each worth a stated number of points.

  4. Recorded judge

    Claude Opus 5 labels each criterion, with its evidence.

  5. Deterministic score

    Points add up. No weighting applied after the fact.

View GitHub Read the capability taxonomy

The judge is a record, not a call The five rules in full, what is guaranteed, and what is not.
  1. Every model receives the same scenario, source order and benchmark prompt.
  2. Every official model is run three times.
  3. Monday Score is the mean of those three official runs.
  4. Speed and cost are reported separately and never affect the quality score.
  5. Every input, raw response, judge verdict and scoring rule is public.

About the judge

The judge is Claude Opus 5 reading each response against the rubric in a recorded session, not an API call. That means a published result is reproducible as a record, since every label ships with its evidence and can be checked line by line, but a third party cannot regenerate the labels from scratch by running a command. The same effort that built the harness also produced the labels, which is a real bias risk rather than a hypothetical one. The audit trail is the mitigation, not a claim that none was needed.

// model fingerprints

Two models, where they differ

The same criteria as the leaderboard, regrouped by what they measure. Showing the 6 capabilities with the widest gap across the field, among those measured by at least 3 criteria: 174 of 500 rubric points.

vs

Qwen 3.8 Flash

95.1 Monday Score

GLM 5.3 Flash

94.9 Monday Score

  1. temporal reasoning 3 crit · 14 pts gap 14
    79%
    64%
  2. supersession 6 crit · 27 pts gap 7
    93%
    100%
  3. causal discipline 8 crit · 23 pts gap 6
    100%
    94%
  4. uncertainty calibration 7 crit · 23 pts gap 5
    89%
    94%
  5. quantitative reasoning 16 crit · 66 pts gap 5
    93%
    98%
  6. business judgment 6 crit · 21 pts gap 0
    98%
    98%

Read a row, not a column: two models are comparable on one capability, two capabilities are not comparable on one model. The point mass is whatever the rubrics gave them, not a claim about importance.

Explore all 22 capabilities

// trust the result

Built to be audited

Auditable is not a claim about how the numbers came out. It is a set of things done before they existed, each of which someone outside this can check.

  1. Nothing is discarded for scoring badly

    Every official run is scored and published, the bad ones included. There is no best-of-N and no retry after a poor result, so the 3 runs behind a score are the 3 runs that happened.

  2. The rubric is frozen before anything runs

    Criteria and their points are fixed and versioned ahead of the first run. A later change opens a new rubric version rather than rewriting the one a published score was measured against.

  3. Every verdict ships with its evidence

    The judge labels one atomic criterion at a time, and each label carries the passage of the model’s own answer it was based on. A disagreement can be taken to the sentence it turns on.

  4. Every model gets the identical input

    Same scenario, same source order, same benchmark prompt. Nothing is tuned per model, and no model sees a running total while it answers.

  5. The release is hashed and checkable from a clean clone

    Each artefact carries a hash and the version carries one digest over all of them. Clone the tag, recompute, and you get the same digest or you have found something.

  • Fresh-clone verified
  • Byte-identical release

Inspect v1.7 on GitHub

Technical release details The suite digest, and what it does and does not cover.

Verified from a clone of the tag rather than from a working copy: all five freeze manifests clean and every group digest matching. The digest below is the whole version in sixty-four characters.

Suite digest

060a2cdd79fcf6096039af5af4b11f4e96fabc683476c63579f01043d398ee45

This is a claim about the evidence trail, not about the judgement. The trail reproduces exactly; the generative labelling does not, and the methodology says so.

MondayBench v1.1 introduces scenario-specific evaluator contract fingerprints. Candidate generations and scoring semantics are unchanged from v1.0; existing recorded judgments were preserved, and every published score from v1.0 is identical here. GLM 5.3 is the first new model published under the v1.1 evaluator contract. v1.0 stays archived and byte-reproducible.

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.