// mondaybench #003 beta

We gave 10 open models
the same messy Monday morning

Six sources. Six dated commitments across one week, three of them waiting on the same three-hour migration, and one engineer claimed by two of them at once.

One question: Can it plan through dependencies?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable

// six dated commitments · one week · 3 waits on the billing migration

Mon Tue Wed Thu Fri Contoso Q3 compliance export Julia Tue 16:00 Enterprise SSO live for Atlas Marcus Wed 17:00 Billing platform Q3 milestone Elena Thu 15:00 Public pricing page Sarah Thu 17:00 Board pack, revenue section Sarah Fri 10:00 Self-serve onboarding launch Julia Fri 12:00
  • now · Monday 09:10
  • waits on the billing migration
  • owes it nothing
01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 GPT-6 Astra gpt-6-astra closed 100.0 0.0 — —
2 Claude Opus 5 closed 98.3 5.0 – –
3 Fable 5.1 closed 96.7 5.0 – –
1 4 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 96.3 2.5 81.4s 89.2s
2 5 DeepSeek V4 Flash deepseek-v4-flash 95.0 0.0 136.0s 143.0s
3 6 GLM 5.3 glm5.3 94.8 5.0 253.3s 286.6s
4 7 Qwen 3.8 Flash qwen3.8-flash 94.5 11.5 309.3s 327.2s
5 8 GLM 5.3 Flash glm5.3-flash 92.8 19.0 120.8s 145.9s
6 9 MIMO v2.5 mimo-v2.5 90.3 11.5 139.1s 167.8s
7 10 GLM 5.2 glm5.2 87.3 14.0 40.5s 64.3s
11 Gemini 3.8 Flash (High) closed 85.3 23.5 – 30.0s
8 12 Qwen 3.6 qwen3.6 65.5 13.0 0.9s 14.8s
9 13 Gemma 4 gemma4 65.3 8.0 0.6s 7.7s
  • Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
  • TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.

// quality vs speed

The smartest model isn’t always the one you want to wait for

GPT-6 Astra tops this scenario at 100.0, after — before its first answer token. GPT-6 Astra starts answering in — and scores 100.0. Models trade quality for response time in very different ways.

60 70 80 90 100 0s100s200s300s TTFA median, seconds to first answer Monday Score GPT-6 Astra 100.0 · — DeepSeek V4.1 Flash 96.3 · 81.4s DeepSeek V4 Flash 95.0 · 136.0s GLM 5.3 94.8 · 253.3s Qwen 3.8 Flash 94.5 · 309.3s GLM 5.3 Flash 92.8 · 120.8s MIMO v2.5 90.3 · 139.1s GLM 5.2 87.3 · 40.5s Qwen 3.6 65.5 · 0.9s Gemma 4 65.3 · 0.6s

Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.

Monday Score and time to first answer, per model
Model Monday Score TTFA median
GPT-6 Astra 100.0 —
DeepSeek V4.1 Flash 96.3 81.4s
DeepSeek V4 Flash 95.0 136.0s
GLM 5.3 94.8 253.3s
Qwen 3.8 Flash 94.5 309.3s
GLM 5.3 Flash 92.8 120.8s
MIMO v2.5 90.3 139.1s
GLM 5.2 87.3 40.5s
Qwen 3.6 65.5 0.9s
Gemma 4 65.3 0.6s

// what stood out

Three things worth saying out loud

Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.

  1. 01

    9 / 24

    runs never got the two chains running at once

    Marcus

    One engineer was claimed by two deadlines. Nine plans queued them behind him

    Marcus is the default owner of three items and the only possible owner of two of them. The third — the Contoso permission patch — has a named substitute: Daniel, who was inside the permission model two weeks ago, at a cost of an hour and a half of his time and half an hour of Marcus’s for the walkthrough. Nine runs put both chains through Marcus one after the other while Daniel had an empty week. The sources state what each of them can do; not one of them states who should do what.

  2. 02

    10 / 24

    runs found both halves of the gate

    Pricing

    A plan that blocks itself, stated in two consecutive sentences

    Sarah intends to assemble the packaging approval pack from screenshots of the finished pricing page — and the page cannot go public until that approval clears. Nobody in the scenario notices, and the way out is a line in Priya’s checklist: design mockups are accepted. Ten runs of twenty-four stated the gate and used the mockup route; three stated neither, and the Thursday page date goes with them.

  3. 03

    81.0 – 100.0

    three runs, same prompt

    GLM 5.3 Flash

    The spread inside one model is wider than the gap across the top five

    Fewer than five points separate first place from fifth. GLM 5.3 Flash’s own three runs are nineteen apart on identical bytes: a perfect 100.0, a 97.5, and an 81.0 that moves the onboarding launch from Friday to Thursday and then queues both chains through one engineer. Qwen 3.8 Flash lost first place the same way — 100.0 and 95.0, then an 88.5. DeepSeek V4 Flash won it by scoring 95.0 three times without moving. Read the range column beside the score, or do not read the score.

02

The test

What they were given, and what they had to work out.

// the test

Every fact is written down
Not one conclusion is

Six ordinary files: Slack, email, meeting notes, a Friday task list and two CSVs. Durations, owners, deadlines and who is capable of what are all stated plainly. What the model is never given is this picture.

What has to happen before what

Billing schema migration Elena · 3h API compatibility patch Marcus · 4h SSO regression QA · 2h Production rollout Marcus · 1h Wed 17:00 Atlas, committed in writing Contoso permission patch Marcus · 2h Q3 compliance export Julia · 1h Tue 16:00 Contoso MSA 8.3, contractual Reconciliation backfill Elena · 1.5h Board pack, revenue Sarah · 2h Fri 10:00 Board meeting Packaging approval pack Sarah · 1.5h Committee approval Public pricing page Sarah · 3h Thu 17:00 Public pricing page live
  • only this person can do it
  • has a named substitute

9 dependencies. Each is stated once, in a Slack thread, an email or a line of the meeting notes, and no source states two of them together.

The sentence no source writes

Marcus

  • API compatibility patch 4h · only this person can do it
  • Production rollout 1h · only this person can do it
  • Contoso permission patch 2h · Daniel 3.5h + 30min

Marcus is the default owner of three items and the only possible owner of two of them. The third has a substitute, and swapping it costs an hour and a half of Daniel’s time plus half an hour of Marcus’s for the walkthrough. Both dated commitments can then be met. Nothing in the material says so.

He is also holding 3 hours for onboarding instrumentation, which Friday’s task list still calls a launch blocker and which Julia released the same evening.

A plan that closes on itself

  1. Sarah plans to build the approval pack from screenshots of the finished pricing page.
  2. The page cannot go public until the packaging committee has approved it.
  3. Priya’s checklist accepts design mockups. Captures from a live page were never required.

What happened

  • 15 of 24 official runs ran the two chains at the same time
  • 17 of 24 official runs resolved the contested engineer in full
  • 10 of 24 official runs found both halves of the pricing gate
03

The method

How the score is built, and how to check it.

// scoring

How Monday Score works

One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.

  1. Critical path and dependencies Does it work out what has to happen before what? 21
  2. Resource allocation Does it keep the one irreplaceable person on the irreplaceable work? 25
  3. Planning and parallelisation Does the plan use the day, or queue it? 18
  4. Decision quality Does it act on the real blocker rather than the loudest one? 14
  5. State reconstruction Does it know what is currently true? 10
  6. Accuracy and uncertainty Does it carry the given values through, and stop where the evidence stops? 12
  7. total 100

The judge does not assign a 0–⁠100 score directly

It classifies atomic criteria one at a time, 21 of them in rubric 0.3, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.

  • PASS full points for the criterion
  • PARTIAL half points
  • FAIL nothing

Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.

Penalties

Applied on top of the dimension score, capped at −25 per run.

  • -5 Material hallucination A claim that matters to the briefing and is not supported by the material.
  • -7 Prohibited action Recommending something the current state explicitly rules out.
  • -7 Major planning error Recommending a plan that cannot run under the stated dependencies or availability.
  • -5 Major contradiction Contradicting itself about the state of an important issue.

see_the_methodology →

// provenance

Exactly what produced these numbers

Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.

Scenario

monday-003

Benchmark version

0.3

Rubric

0.3

Evaluator

evaluator-opus5-v0.3

Judge

claude-opus-5

Result generated

2026-09-11

Prompt SHA-256

aaf788f309d955d2b427c86fe7d95d2bbf88575fde7306adcfd74b0ec4a4103f

Attempts needed for three valid runs

A run can be rejected before it is ever scored: truncated, or returning no reasoning when reasoning was requested, or a cached duplicate. Those never reach a leaderboard, because a leaderboard can only exclude what it evaluated. This is a property of the model and its endpoint rather than of the answer, and it is reported apart from the score.

  • Qwen 3.8 Flash 3 of 7 attempts REASONING_NOT_DELIVERED

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.