// mondaybench · v1.5

Can a model handle
your Monday?

Five scenarios rebuild one real working morning. Every model gets the same inbox, the same contradictions and the same half-finished decisions, answers 3 times, and is scored against a frozen rubric.

// overall

Overall leaderboard

Monday Score is the equal-weight average across all five real-work scenarios.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

Leaderboard

9 open-weight models, ranked by Monday Score. 13 models ranked by Monday Score, including the 4 closed-weight models.

# Model Monday Score Δ top Worst Monday Scenario spread
01 closed
98.80
+3.73 96.0
4.0
02 closed
97.07
+2.00 −1.73 90.8
8.8
01 03
95.07
−3.73 92.5
4.5
02 04
94.93
−0.14 −3.87 87.3
11.2
05 closed
94.90
−0.17 −3.90 89.3
10.7
03 06
94.53
−0.54 −4.27 90.0
8.3
04 07
93.73
−1.34 −5.07 87.7
9.8
05 08
91.70
−3.37 −7.10 86.0
11.3
09 closed
88.37
−6.70 −10.43 78.5
21.5
06 10
86.67
−8.40 −12.13 78.5
16.7
07 11
86.50
−8.57 −12.30 71.2
27.3
08 12
68.77
−26.30 −30.03 46.0
36.8
09 13
63.00
−32.07 −35.80 35.2
54.5
Expand scenario scores The five Mondays behind each average, shaded by score.

Cell shade 35.2 100.0

Model #001 #002 #003 #004 #005 Scenario spread
closed 100.0 100.0 100.0 98.0 96.0 4.0
closed 98.5 98.0 98.3 90.8 99.7 8.8
92.5 94.8 94.5 96.5 97.0 4.5
98.0 87.3 92.8 98.0 98.5 11.2
closed 97.5 91.0 96.7 89.3 100.0 10.7
95.0 93.0 96.3 90.0 98.3 8.3
97.5 87.7 94.8 91.7 97.0 9.8
90.8 89.3 95.0 86.0 97.3 11.3
closed 97.0 81.0 85.3 78.5 100.0 21.5
88.2 81.2 90.3 78.5 95.2 16.7
84.3 91.2 87.3 71.2 98.5 27.3
82.8 70.2 65.5 46.0 79.3 36.8
70.7 54.2 65.3 35.2 89.7 54.5
Technical details Run stability, unranked models, and what produced each number.

Run stability

How far apart the three official runs of one scenario land, averaged across the five. A model you cannot get the same answer from twice is a different product from one you can, so it is published here and never counted in the Monday Score.

Model Mean run range Worst run range
closed 0.3 1.5
closed 2.5 8.5
closed 3.5 9.5
4.3 10.5
closed 5.4 23.5
6.0 19.0
6.4 11.5
6.8 13.5
7.0 17.0
9.1 16.0
9.9 19.0
14.7 31.0
16.3 33.0

// what it costs

Three models the benchmark cannot separate

The top three scores sit inside the measurement noise, so the ranking stops being the interesting question. What separates them is what they charge, how long they take and how much they write.

Top performance band

  • Highest in band

    Qwen 3.8 Flash

    95.07 / 100

    $12.61–$15.56 / 1k · 258.9s

  • Cheapest top-band

    GLM 5.3 Flash

    94.93 / 100

    $10.71 / 1k

  • Most concise top tier

    Claude Fable 5.1

    94.90 / 100

    50.9s · ~8.2k output

Fable reaches the same measured quality band roughly 5.08×–5.46× faster while generating 2.45×–3.87× fewer output tokens.

The trade-off: 29.3×–42.6× the cost.

Result views

Quality vs cost

Opus records the highest measured quality, but at a very different price point from the leading open models. 97.07 at $246.91 per 1,000 tasks against 94.93 at $10.71.

GLM 5.3 appears twice: $99.20 per 1,000 tasks at the vendor’s list price, $5.67 on the metered route the benchmark actually ran. Same model, same score, 17.5× apart — model price and route price can tell very different stories.

Provider view mixes first-party list prices with public third-party provider prices.

60 70 80 90 100 $0.10 $1 $10 $100 $1000 Cost per 1,000 MondayBench-equivalent tasks · Log scale Claude Opus 5 — 97.07 · $246.91/1k · First-party · HIGH Claude Opus 5 Qwen 3.8 Flash — 95.07 · $12.61–$15.56/1k · First-party · MEDIUM Qwen 3.8 Flash GLM 5.3 Flash — 94.93 · $10.71/1k · First-party · HIGH GLM 5.3 Flash Claude Fable 5.1 — 94.90 · $456.27/1k · First-party · HIGH Claude Fable 5.1 DeepSeek V4.1 Flash — 94.53 · $10.72–$22.47/1k · First-party · MEDIUM DeepSeek V4.1 Flash GLM 5.3 — 93.73 · $99.20/1k · First-party · HIGH GLM 5.3 DeepSeek V4 Flash — 91.70 · $12.94–$25.88/1k · First-party · HIGH DeepSeek V4 Flash MIMO v2.5 — 86.67 · $3.33/1k · OpenRouter · HIGH MIMO v2.5 GLM 5.2 — 86.50 · $39.50/1k · First-party · HIGH GLM 5.2 Qwen 3.6 — 68.77 · $7.70/1k · OpenRouter · MEDIUM Qwen 3.6 Gemma 4 — 63.00 · $0.44/1k · OpenRouter · MEDIUM Gemma 4 GLM 5.3 — 93.73 · $5.67/1k · Helmcode route GLM 5.3 · route
First-party only OpenRouter Helmcode route Best quality for the price · vendor list Best quality for the price · route

From $0.44 to $456 per 1,000 tasks. More than 1037× difference in price across the field. A pricing range only — not a quality comparison.

Quality vs cost
Model Monday Score Cost / 1k Price basis Confidence
Claude Opus 5 97.07 $246.91 First-party HIGH
Qwen 3.8 Flash On frontier 95.07 $12.61–$15.56 First-party MEDIUM
GLM 5.3 Flash On frontier 94.93 $10.71 · promo $5 First-party HIGH
Claude Fable 5.1 94.90 $456.27 First-party HIGH
DeepSeek V4.1 Flash 94.53 $10.72–$22.47 First-party MEDIUM
GLM 5.3 93.73 $99.20 $5.67 · route First-party identity contested Helmcode route HIGH
DeepSeek V4 Flash 91.70 $12.94–$25.88 First-party HIGH
Gemini 3.8 Flash (High) 88.37 Unknown
MIMO v2.5 86.67 $3.33 OpenRouter HIGH
GLM 5.2 86.50 $39.50 First-party HIGH
Qwen 3.6 68.77 $7.70 OpenRouter ref MEDIUM
Gemma 4 63.00 $0.44 OpenRouter ref MEDIUM
Cost at scale
Model 1k 100k 1M
Qwen 3.8 Flash $13–$16$1.3k–$1.6k$12.6k–$15.6k
GLM 5.3 Flash $11$1.1k$10.7k
Claude Fable 5.1 $456$45.6k$456k
Claude Opus 5 $247$24.7k$247k
Claude Fable 5.1 vs Qwen 3.8 Flash · 29.3×–36.2× +$444+$44.4k+$444k
Claude Fable 5.1 vs GLM 5.3 Flash · 42.6× +$446+$44.6k+$446k

MondayBench-equivalent tasks. A linear extrapolation of the observed workload, not a universal per-request cost.

How each model gets to its result

Opus combines the highest measured score with the fastest first answer, but at a much higher cost than the leading open models.

Two observable dimensions. Neither is folded into Monday Score, and there is no efficiency score.

Speed

60 70 80 90 100 0s60s120s180s240s300s Median time to first answer Qwen 3.8 Flash — 95.07 · 258.9s · ~31.7k output Qwen 3.8 Flash GLM 5.3 Flash — 94.93 · 278.2s · ~20.1k output GLM 5.3 Flash Claude Fable 5.1 — 94.90 · 50.9s · ~8.2k output Claude Fable 5.1 DeepSeek V4.1 Flash — 94.53 · 69.2s · ~17.5k output DeepSeek V4.1 Flash GLM 5.3 — 93.73 · 253.3s · ~21.3k output GLM 5.3 DeepSeek V4 Flash — 91.70 · 121.4s · ~18.2k output DeepSeek V4 Flash MIMO v2.5 — 86.67 · 170.4s · ~13.1k output MIMO v2.5 GLM 5.2 — 86.50 · 44.1s · ~8.1k output GLM 5.2 Qwen 3.6 — 68.77 · 1.0s · ~3.3k output Qwen 3.6 Gemma 4 — 63.00 · 0.6s · ~0.9k output Gemma 4 fastest measured Claude Opus 5 — 97.07 · 41.3s · ~9.2k output Claude Opus 5

Fable reaches the same measured quality band roughly 5.08×–5.46× faster.

Concision

  • Gemma 4 ~0.9k Provider-reported reasoning tokens: ~0.0k
  • Qwen 3.6 ~3.3k Provider-reported reasoning tokens: ~1.2k
  • GLM 5.2 ~8.1k Provider-reported reasoning tokens: ~5.9k
  • Claude Fable 5.1 ~8.2k Provider-reported reasoning tokens: ~4.4k
  • Claude Opus 5 ~9.2k Provider-reported reasoning tokens: ~4.8k
  • MIMO v2.5 ~13.1k Provider-reported reasoning tokens: ~0.0k
  • DeepSeek V4.1 Flash ~17.5k Provider-reported reasoning tokens: ~15.6k
  • DeepSeek V4 Flash ~18.2k Provider-reported reasoning tokens: ~16.5k
  • GLM 5.3 Flash ~20.1k Provider-reported reasoning tokens: ~18.0k
  • GLM 5.3 ~21.3k Provider-reported reasoning tokens: ~19.3k
  • Qwen 3.8 Flash ~31.7k Provider-reported reasoning tokens: ~27.4k

Average output tokens

Fable produces substantially shorter outputs on this workload.

2.45× fewer output tokens than GLM 5.3 Flash and 3.87× fewer than Qwen 3.8 Flash.

Observable generated output only. Provider observability is not homogeneous across models — detail only, not a comparison.

Not plotted: Gemini 3.8 Flash (High). The route reports no time-to-first-answer and no token count, and neither is estimated. Total generation latency for those runs is in the per-scenario tables.

// what the table hides

Three things the ranking does not say

  1. 4 / 5

    scenarios won outright by the model ranked first

    The winner only wins one scenario

    GPT-6 Astra ranks #1 because it never has a bad Monday.

    Worst scenario: 96.0

  2. 98.8 ≠ 95.1

    two scores, two different machines

    Almost the same score. Very different models

    GPT-6 has the higher floor. Qwen reaches higher peaks.

    Range across scenarios: 4.0 vs 4.5

  3. 6 / 10

    models score their worst on #004

    Business data is where models break

    Numbers Don’t Lie is the worst-scoring scenario for 6 of 10 models.

    Widest field on the board

// the finding

Models can keep a secret
They don’t always know who needs it

0 / 30

explicit external leaks

11 / 30

failed to route the security item to engineering

The same sensitive fact had to stay away from the customer and the public status page, while still reaching Engineering.

See Read the Room

// model fingerprints

Two models, where they differ

The same criteria as the leaderboard, regrouped by what they measure. Showing the 6 capabilities with the widest gap across the field, among those measured by at least 3 criteria: 174 of 500 rubric points.

vs

Qwen 3.8 Flash

95.1 Monday Score

GLM 5.3 Flash

94.9 Monday Score

  1. temporal reasoning 3 crit · 14 pts gap 14
    79%
    64%
  2. supersession 6 crit · 27 pts gap 7
    93%
    100%
  3. causal discipline 8 crit · 23 pts gap 6
    100%
    94%
  4. uncertainty calibration 7 crit · 23 pts gap 5
    89%
    94%
  5. quantitative reasoning 16 crit · 66 pts gap 5
    93%
    98%
  6. business judgment 6 crit · 21 pts gap 0
    98%
    98%

Read a row, not a column: two models are comparable on one capability, two capabilities are not comparable on one model. The point mass is whatever the rubrics gave them, not a claim about importance.

Explore all 22 capabilities

// method

How MondayBench works

Every input, response, verdict and scoring rule is public.

  1. Same input

    Identical scenario, source order and prompt for every model.

  2. 3 runs

    Every official model answers three times.

  3. Atomic rubric

    Frozen criteria, each worth a stated number of points.

  4. Recorded judge

    Claude Opus 5 labels each criterion, with its evidence.

  5. Deterministic score

    Points add up. No weighting applied after the fact.

View GitHub Read the capability taxonomy

The judge is a record, not a call The five rules in full, what is guaranteed, and what is not.
  1. Every model receives the same scenario, source order and benchmark prompt.
  2. Every official model is run three times.
  3. Monday Score is the mean of those three official runs.
  4. Speed and cost are reported separately and never affect the quality score.
  5. Every input, raw response, judge verdict and scoring rule is public.

About the judge

The judge is Claude Opus 5 reading each response against the rubric in a recorded session, not an API call. That means a published result is reproducible as a record, since every label ships with its evidence and can be checked line by line, but a third party cannot regenerate the labels from scratch by running a command. The same effort that built the harness also produced the labels, which is a real bias risk rather than a hypothetical one. The audit trail is the mitigation, not a claim that none was needed.

// trust the result

Built to be audited

Auditable is not a claim about how the numbers came out. It is a set of things done before they existed, each of which someone outside this can check.

  1. Nothing is discarded for scoring badly

    Every official run is scored and published, the bad ones included. There is no best-of-N and no retry after a poor result, so the 3 runs behind a score are the 3 runs that happened.

  2. The rubric is frozen before anything runs

    Criteria and their points are fixed and versioned ahead of the first run. A later change opens a new rubric version rather than rewriting the one a published score was measured against.

  3. Every verdict ships with its evidence

    The judge labels one atomic criterion at a time, and each label carries the passage of the model’s own answer it was based on. A disagreement can be taken to the sentence it turns on.

  4. Every model gets the identical input

    Same scenario, same source order, same benchmark prompt. Nothing is tuned per model, and no model sees a running total while it answers.

  5. The release is hashed and checkable from a clean clone

    Each artefact carries a hash and the version carries one digest over all of them. Clone the tag, recompute, and you get the same digest or you have found something.

  • Fresh-clone verified
  • Byte-identical release

Inspect v1.5 on GitHub

Technical release details The suite digest, and what it does and does not cover.

Verified from a clone of the tag rather than from a working copy: all five freeze manifests clean and every group digest matching. The digest below is the whole version in sixty-four characters.

Suite digest

40c55b6c4b7cf5c597ee81edc5b26d416c1db92e949bb817766a2aeb5c9e3ba3

This is a claim about the evidence trail, not about the judgement. The trail reproduces exactly; the generative labelling does not, and the methodology says so.

MondayBench v1.1 introduces scenario-specific evaluator contract fingerprints. Candidate generations and scoring semantics are unchanged from v1.0; existing recorded judgments were preserved, and every published score from v1.0 is identical here. GLM 5.3 is the first new model published under the v1.1 evaluator contract. v1.0 stays archived and byte-reproducible.

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.