// mondaybench #002 beta

We gave 10 open models
the same messy Monday morning

Nine sources. A release that failed on Friday, two ways to recover it, and four people whose calendars barely overlap.

One question: Can it tell what is still true?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 GPT-6 Astra gpt-6-astra closed 100.0 0.0 — —
2 Claude Opus 5 closed 98.0 3.0 – –
1 3 Qwen 3.8 Flash qwen3.8-flash 94.8 8.5 258.9s 285.0s
2 4 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 93.0 10.5 61.2s 67.8s
3 5 GLM 5.2 glm5.2 91.2 8.0 67.0s 85.4s
6 Fable 5.1 closed 91.0 6.0 – –
4 7 DeepSeek V4 Flash deepseek-v4-flash 89.3 13.5 162.6s 178.8s
5 8 GLM 5.3 glm5.3 87.7 4.0 176.2s 197.1s
6 9 GLM 5.3 Flash glm5.3-flash 87.3 6.5 238.5s 316.6s
7 10 MIMO v2.5 mimo-v2.5 81.2 33.0 163.3s 195.4s
11 Gemini 3.8 Flash (High) closed 81.0 0.0 – 38.8s
8 12 Qwen 3.6 qwen3.6 70.2 31.0 0.9s 18.2s
9 13 Gemma 4 gemma4 54.2 13.5 0.3s 8.7s
  • Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
  • TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.

// quality vs speed

The smartest model isn’t always the one you want to wait for

GPT-6 Astra tops this scenario at 100.0, after — before its first answer token. GPT-6 Astra starts answering in — and scores 100.0. Models trade quality for response time in very different ways.

50 60 70 80 90 100 0s100s200s TTFA median, seconds to first answer Monday Score GPT-6 Astra 100.0 · — Qwen 3.8 Flash 94.8 · 258.9s DeepSeek V4.1 Flash 93.0 · 61.2s GLM 5.2 91.2 · 67.0s DeepSeek V4 Flash 89.3 · 162.6s GLM 5.3 87.7 · 176.2s GLM 5.3 Flash 87.3 · 238.5s MIMO v2.5 81.2 · 163.3s Qwen 3.6 70.2 · 0.9s Gemma 4 54.2 · 0.3s

Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.

Monday Score and time to first answer, per model
Model Monday Score TTFA median
GPT-6 Astra 100.0 —
Qwen 3.8 Flash 94.8 258.9s
DeepSeek V4.1 Flash 93.0 61.2s
GLM 5.2 91.2 67.0s
DeepSeek V4 Flash 89.3 162.6s
GLM 5.3 87.7 176.2s
GLM 5.3 Flash 87.3 238.5s
MIMO v2.5 81.2 163.3s
Qwen 3.6 70.2 0.9s
Gemma 4 54.2 0.3s

// what stood out

Four things worth saying out loud

Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.

  1. 01

    4 / 24

    runs chose the safer path

    Option B

    Almost every model traded safety for a deadline it was not short of

    Both recovery paths open rollout at 11:00, an hour inside the cutoff. Option B carries an eighth of the session-reset risk at no cost in time. Twenty of twenty-four runs decided it did not fit — most of them by starting the sequence at 09:15, when the engineer arrives, rather than at 09:07, when the diff can be submitted.

  2. 02

    #1 → #5

    between #001 and #002

    GLM 5.3 Flash

    The model that won the first scenario came fifth in this one

    Not a regression — a different exam. #001 rewards noticing what changed; #002 rewards building a plan that runs. GLM 5.3 Flash’s three runs land within 6.5 points of each other, and not one of them chose the safer release path.

  3. 03

    64.5 – 97.5

    three runs, same prompt

    MIMO v2.5

    One model produced its best and its worst answer on the same input

    A 33-point range, the widest here. Its best run is the second-best answer in the whole battery; its worst states that Security verification is not needed and then schedules it anyway. Three runs per model exists for exactly this: on one run, MIMO looks like a top-two model or a bottom-three one depending on which you draw.

  4. 04

    3 / 3

    plans that cannot be executed

    Gemma 4

    Being wrong is ordinary. Recommending an impossible plan is not

    All three of its runs schedule the deploy to start after the only engineer who can run it has left — one of them at 10:30, ten minutes past his hard stop. That is what the MAJOR_PLANNING_ERROR penalty is scoped to, and Gemma 4 is the only model that triggered it on every one of its runs.

02

The test

What they were given, and what they had to work out.

// the test

Everything it needed was stated
None of it was stated together

A release failed on Friday. Two ways to recover it, four people whose calendars barely overlap, and a customer window that closes at noon. No fact here is hidden: the work is fitting them together.

Who is free, and when

09:00 09:30 10:00 10:30 11:00 11:30 12:00 Marcus Priya Sam Daniel
  • now · 09:07
  • Marcus leaves · 10:20
  • Atlas rollout must have begun · 12:00

The two recovery paths · A

2.9.4

Retry the build that already failed once, with a shard precheck.

deploy
20 min
session-reset risk
8%
 
no code change

The two recovery paths · B

2.9.5-patch

A patch that avoids the migration path that broke. Needs a Security delta review first.

deploy
45 min
session-reset risk
<1%
 
code change

Option B, sequenced

09:00 09:30 10:00 10:30 11:00 11:30 12:00 Priya Security delta review · 25m Marcus Deploy 2.9.5-patch · 45m Priya Post-deploy verification · 20m Sam Open and sanity-check rollout · 15m

The deploy finishes 3 minutes before Marcus leaves. That is the whole margin.

The step that makes it fit

Priya reviews asynchronously from a submitted diff, and she is free from 09:00. The chain starts at 09:07 with an action the user can take, not at 09:15 when Marcus arrives.

What happened

Both paths open rollout at 11:00, 60 minutes inside the deadline. Option B carries an eighth of the risk at no cost in time. 4 of 24 official runs chose it; 20 of 24 never built a schedule that showed it fits.

Miss it and migration slips to Thursday. The renewal is signed, so the cost is executive noise rather than revenue.

03

The method

How the score is built, and how to check it.

// scoring

How Monday Score works

One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.

  1. State reconstruction Does it know what is currently true? 15
  2. Temporal supersession Does it notice what a later message overruled? 15
  3. Planning and dependencies Can it build a plan that would actually run? 25
  4. Decision and prioritisation Does it act on the tightest constraint first? 20
  5. Cross-source reasoning Does it combine files, and keep arithmetic apart from cause? 15
  6. Uncertainty Does it refuse to settle what the evidence cannot? 10
  7. total 100

The judge does not assign a 0–⁠100 score directly

It classifies atomic criteria one at a time, 20 of them in rubric 0.2, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.

  • PASS full points for the criterion
  • PARTIAL half points
  • FAIL nothing

Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.

Penalties

Applied on top of the dimension score, capped at −30 per run.

  • -5 Material hallucination A claim that matters to the briefing and is not supported by the material.
  • -10 Prohibited action Recommending something the current state explicitly rules out.
  • -7 Major planning error Recommending a plan that cannot run under the stated dependencies or availability.
  • -5 Major contradiction Contradicting itself about the state of an important issue.

see_the_methodology →

// provenance

Exactly what produced these numbers

Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.

Scenario

monday-002

Benchmark version

0.2

Rubric

0.2

Evaluator

evaluator-opus5-v0.2

Judge

claude-opus-5

Result generated

2026-09-11

Prompt SHA-256

740e627cbc7413c321ede0d54000504de54821d34ff4677f82731b89600423a8

Attempts needed for three valid runs

A run can be rejected before it is ever scored: truncated, or returning no reasoning when reasoning was requested, or a cached duplicate. Those never reach a leaderboard, because a leaderboard can only exclude what it evaluated. This is a property of the model and its endpoint rather than of the answer, and it is reported apart from the score.

  • Qwen 3.8 Flash 3 of 4 attempts REASONING_NOT_DELIVERED

One run was generated and scored but does not count

  • monday-002__qwen3.8-flash__002 66.5 The provider served this run without the reasoning phase its two sibling runs used — zero reasoning tokens against 42,808 and 30,342 — so it was not comparable with them. Its score stands and its evidence is published; it simply does not count.

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.