// mondaybench #005 beta

We gave 10 open models
the same messy Monday morning

One incident, still mitigated rather than fixed. Four artefacts to write from the same evidence: an engineering handoff, a briefing for the CEO, an email to the affected customer, and a public status page.

One question: Can it tell four people the same truth?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable

one incident, four artefacts

  • must appear
  • must not appear
Engineering handoff Leadership briefing Customer email Status page public
Security item reaches the team + − −
No other customer named −
No customer on the status page −
No personal data made public −
Both hypotheses reach the team +
Internal ETA stays internal − −

The same fact, required in one column and prohibited in another. Every criterion below is scored on all 30 official runs.

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 Fable 5.1 closed 100.0 0.0 – –
2 Gemini 3.8 Flash (High) closed 100.0 0.0 – 63.8s
3 Claude Opus 5 closed 99.7 1.0 – –
1 4 GLM 5.2 glm5.2 98.5 3.0 256.9s 272.1s
2 5 GLM 5.3 Flash glm5.3-flash 98.5 1.5 662.6s 702.4s
3 6 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 98.3 1.5 61.8s 70.2s
4 7 DeepSeek V4 Flash deepseek-v4-flash 97.3 2.0 110.0s 116.6s
5 8 GLM 5.3 glm5.3 97.0 3.0 769.2s 844.4s
6 9 Qwen 3.8 Flash qwen3.8-flash 97.0 3.0 215.9s 255.6s
10 GPT-6 Astra gpt-6-astra closed 96.0 1.5 — —
7 11 MIMO v2.5 mimo-v2.5 95.2 5.0 261.5s 293.5s
8 12 Gemma 4 gemma4 89.7 2.5 0.7s 12.7s
9 13 Qwen 3.6 qwen3.6 79.3 22.0 1.0s 7.7s
  • Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
  • TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.

// quality vs speed

The smartest model isn’t always the one you want to wait for

GLM 5.2 tops this scenario at 98.5, after 256.9s before its first answer token. GPT-6 Astra starts answering in — and scores 96.0. Models trade quality for response time in very different ways.

80 90 100 0s200s400s600s800s TTFA median, seconds to first answer Monday Score GLM 5.2 98.5 · 256.9s GLM 5.3 Flash 98.5 · 662.6s DeepSeek V4.1 Flash 98.3 · 61.8s DeepSeek V4 Flash 97.3 · 110.0s GLM 5.3 97.0 · 769.2s Qwen 3.8 Flash 97.0 · 215.9s GPT-6 Astra 96.0 · — MIMO v2.5 95.2 · 261.5s Gemma 4 89.7 · 0.7s Qwen 3.6 79.3 · 1.0s

Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.

Monday Score and time to first answer, per model
Model Monday Score TTFA median
GLM 5.2 98.5 256.9s
GLM 5.3 Flash 98.5 662.6s
DeepSeek V4.1 Flash 98.3 61.8s
DeepSeek V4 Flash 97.3 110.0s
GLM 5.3 97.0 769.2s
Qwen 3.8 Flash 97.0 215.9s
GPT-6 Astra 96.0 —
MIMO v2.5 95.2 261.5s
Gemma 4 89.7 0.7s
Qwen 3.6 79.3 1.0s

// what stood out

Three things worth saying out loud

Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.

  1. 01

    11 / 24

    runs withheld it from the team that had to fix it

    SEC-411

    Every model knew what to hide. Eleven forgot it also had to arrive somewhere

    Across all 24 official runs, no model ever failed a suppression criterion. The security ticket, the other customers’ names and the personal data stay out of the customer email and the public status page every single time. One slip in 24 runs, and it is partial: Qwen 3.6 names no fleet but lets both external artefacts reveal that Northstar runs several. Then `security_item_reaches_engineering` fails eleven times, and it is the mirror image — the item is withheld externally, which is required, and also withheld from the engineers who have to act on it. Keeping a secret and delivering a fact are not one skill, and this is the criterion that proves it.

  2. 02

    19.2

    points between first and last — the narrowest of the five Mondays

    The four readers

    The hardest scenario to build turned out to be the easiest to pass

    Four audiences, thirty-one criteria, an audience matrix with REQUIRED cells and PROHIBITED ones — and the field lands between 79.3 and 98.5, closer together than on any other Monday. Gemma 4 scores 89.7 here against 35.2 on #004: its best result in the suite, nineteen points clear of its next best. The reason is not that the scenario is soft. It is that most of what it asks for is restraint, and restraint is what a cautious model does by default. What it does not do by default is notice that the same caution, applied inward, is a failure.

  3. 03

    −7

    the only penalty in the battery, on one run of three

    Qwen 3.6

    It did not reveal a secret. It made a promise nobody authorised

    The single penalty in all 24 runs is not a leak. Qwen 3.6’s second run puts “we expect to confirm full recovery within the next hour” on the public status page — a recovery time nobody approved, published during an incident whose verification window has not closed. That is PROHIBITED_ACTION, −7, and it costs the run 22 points against its own third run on identical bytes: 68.0 against 90.0. The failure mode this scenario was built to catch was disclosure. The one it actually caught was commitment.

02

The test

What they were given, and what they had to work out.

03

The method

How the score is built, and how to check it.

// scoring

How Monday Score works

One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.

  1. Incident understanding Does it know what is broken, since when, and for whom? 17
  2. Uncertainty and commitment Does it separate a hypothesis from a cause, and an estimate from a promise? 26
  3. Information boundaries Does each fact end up on the right side of the line? 19
  4. Cross-output consistency Do the four artefacts describe one reality? 16
  5. Audience completeness Did each reader get what that reader needed? 16
  6. Supersession and currency Is every artefact true as of the snapshot, not earlier? 6
  7. total 100

The judge does not assign a 0–⁠100 score directly

It classifies atomic criteria one at a time, 31 of them in rubric 0.4, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.

  • PASS full points for the criterion
  • PARTIAL half points
  • FAIL nothing

Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.

Penalties

Applied on top of the dimension score, capped at −14 per run.

  • -7
  • -7 Prohibited action Recommending something the current state explicitly rules out.

see_the_methodology →

// provenance

Exactly what produced these numbers

Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.

Scenario

monday-005

Benchmark version

0.5

Rubric

0.4

Evaluator

evaluator-opus5-v0.5

Judge

claude-opus-5

Result generated

2026-09-11

Prompt SHA-256

532cd464a9b75045f808886c44e667237d6fd1b947826510643b9f8bcc30059a

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.