// mondaybench #001 beta

We gave 10 open models
the same messy Monday morning

Slack messages. Emails. Meeting notes. A stale task list. Customer metrics.

One question: Can it figure out what actually matters?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
you log off wed 17:30 you’re back mon 09:07 thursday friday Release moves to Wednesday Launch email leaves your plate Contoso needs confirmation by 11:00
48 Slack messages · 7 emails · 2 meetings three of them change what is already true

metrics.csv Umbrella’s anomaly — in the metrics, mentioned by nobody

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 GPT-6 Astra gpt-6-astra closed 100.0 0.0 — —
2 Claude Opus 5 closed 98.5 0.0 – –
1 3 GLM 5.3 Flash glm5.3-flash 98.0 1.5 118.8s 132.4s
4 Fable 5.1 closed 97.5 1.5 – –
2 5 GLM 5.3 glm5.3 97.5 6.0 181.6s 202.8s
6 Gemini 3.8 Flash (High) closed 97.0 0.0 – 13.2s
3 7 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 95.0 6.0 39.3s 44.1s
4 8 Qwen 3.8 Flash qwen3.8-flash 92.5 6.0 72.2s 92.0s
5 9 DeepSeek V4 Flash deepseek-v4-flash 90.8 8.0 55.4s 63.3s
6 10 MIMO v2.5 mimo-v2.5 88.2 6.5 42.4s 64.1s
7 11 GLM 5.2 glm5.2 84.3 16.0 2.7s 6.9s
8 12 Qwen 3.6 qwen3.6 82.8 5.0 0.8s 6.2s
9 13 Gemma 4 gemma4 70.7 6.5 0.6s 5.5s
  • Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
  • TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.

// quality vs speed

The smartest model isn’t always the one you want to wait for

GPT-6 Astra tops this scenario at 100.0, after — before its first answer token. GPT-6 Astra starts answering in — and scores 100.0. Models trade quality for response time in very different ways.

70 80 90 100 0s50s100s150s TTFA median, seconds to first answer Monday Score GPT-6 Astra 100.0 · — GLM 5.3 Flash 98.0 · 118.8s GLM 5.3 97.5 · 181.6s DeepSeek V4.1 Flash 95.0 · 39.3s Qwen 3.8 Flash 92.5 · 72.2s DeepSeek V4 Flash 90.8 · 55.4s MIMO v2.5 88.2 · 42.4s GLM 5.2 84.3 · 2.7s Qwen 3.6 82.8 · 0.8s Gemma 4 70.7 · 0.6s

Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.

Monday Score and time to first answer, per model
Model Monday Score TTFA median
GPT-6 Astra 100.0 —
GLM 5.3 Flash 98.0 118.8s
GLM 5.3 97.5 181.6s
DeepSeek V4.1 Flash 95.0 39.3s
Qwen 3.8 Flash 92.5 72.2s
DeepSeek V4 Flash 90.8 55.4s
MIMO v2.5 88.2 42.4s
GLM 5.2 84.3 2.7s
Qwen 3.6 82.8 0.8s
Gemma 4 70.7 0.6s

// what actually happened

Four things the results showed

Eight models, twenty-four runs, one judge. These are the patterns worth reporting, not the ones that make the best headline.

GLM 5.3 Flash almost solved Monday

It was not only the highest-scoring model. It was also the most consistent: three runs inside a point and a half of each other, with full marks on action detection, state reconstruction and prioritization.

Where it lost points It spotted the Umbrella anomaly but never fully reasoned about errors relative to request volume.

98.0/100

GLM 5.3 Flash

  • run 198.5
  • run 297.0
  • run 398.5
  • range1.5

Perfect on three of six dimensions

One run isn’t enough

A single run could make GLM 5.2 look like a 75-point model or a 91-point model. That is why every official MondayBench result uses three runs, and why the leaderboard reports the spread next to the score instead of hiding it inside a mean.

DeepSeek V4 Flash

run 1: 95.5 run 2: 87.5 run 3: 89.5

95.5 / 87.5 / 89.5 range 8.0

GLM 5.2

run 1: 87.0 run 2: 75.0 run 3: 91.0

87.0 / 75.0 / 91.0 range 16.0

Some models read the conversation but miss the business

Gemma understood much of the conversation, but missed the most important signal hidden in the customer metrics. Umbrella was never mentioned in any of its three runs: the metrics file was effectively ignored.

Data reasoning State reconstruction stayed high: it followed the thread, it just never opened the spreadsheet.

3/15

Data reasoning

  • Gemma 470.7
  • runs naming Umbrella0/3

Fast doesn’t mean safe

Qwen 3.6 was one of the fastest models in the benchmark, but it was also the only model to receive material hallucination penalties.

The scenario states that proposed pricing stays confidential until it is approved and publicly announced. In two of three runs Qwen 3.6 rewrote that as confidentiality ending after Wednesday’s review, or as an announcement being imminent, and put it in the reply it told the user to send the customer. Its third run states the rule correctly.

2

penalties

  • ttfa0.8s
  • total6.2s
  • score82.8
02

The test

What they were given, and what they had to work out.

// the input

This isn’t a trivia test
It’s a state reconstruction problem

Real work is messy. Information arrives through different channels, tasks become stale, decisions change and the most important signal may never be explicitly flagged.

slack

#incidents Thu 09:48 · Marcus Webb

Rollback complete. Error rates are returning to normal.

#sales Fri 08:51 · Noah Williams

Contoso update: call went really well. Commercial terms are basically agreed.

#product Fri 09:36 · Sam Rivera

The export regression is more annoying than we thought. I’m proposing Wednesday instead of Monday.

#leadership Fri 11:42 · Daniel Foster

The dashboard has it moving from 3.1% to 4.4%. Would be good to understand what’s driving that before the board meeting.

tasks
PriorityTaskOwnerDue
Urgent Investigate Acme API failures stale You Thursday
High Confirm Contoso renewal next steps You Monday
High Prepare Analytics v2 launch email stale You Monday
Medium Analyse enterprise churn You Tuesday
metrics · last 24h
CustomerRequestsErrorsp957d
Acme 482,140 8 312ms +3%
Contoso 821,334 22 284ms +7%
Umbrella 94,217 163 1,821ms −41%
Stark Industries 1,482,012 74 298ms +9%

Everything the models saw, in the order they saw it. The full scenario is frozen and published with the results.

// the answer key

What the model needed to figure out

Five judgements decide most of the score. They are published with the scenario, so you can decide for yourself whether the benchmark is asking the right things.

  1. 01

    Contoso

    • €72k ARR renewal.
    • Legal blocker.
    • Written technical confirmation required before 11:00.

    Verify the backup-deletion behavior with the technical owner and respond before 11:00.

  2. 02

    Acme

    • Task list still says urgent.
    • The incident is already resolved.

    Do not treat it as an active incident.

  3. 03

    Analytics v2

    • Task list says Monday launch.
    • Latest information says Wednesday.
    • The launch email is owned by Julia.

    Don’t execute the stale launch tasks.

  4. 04

    Umbrella

    • Nobody mentions it in Slack or email.
    • The anomaly only exists in the metrics.

    Investigate it today, but don’t invent a cause.

  5. 05

    Churn

    • Enterprise churn rose from 3.1% to 4.4%.
    • There are known cancellations but insufficient evidence to explain the full increase.

    Investigate the breakdown before Tuesday’s board meeting without claiming causality.

03

The method

How the score is built, and how to check it.

// scoring

How Monday Score works

One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.

  1. Action detection Did it find the work that had to happen today? 25
  2. State reconstruction Does it know what is currently true? 25
  3. Prioritization Is the order right, and is nothing stale promoted? 15
  4. Data reasoning Did it read the numbers, and read them correctly? 15
  5. Uncertainty Does it separate what it knows from what it assumes? 10
  6. Noise resistance Does it stay away from work that is already done? 10
  7. total 100

The judge does not assign a 0–⁠100 score directly

It classifies atomic criteria one at a time, 19 of them in rubric 0.2, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.

  • PASS full points for the criterion
  • PARTIAL half points
  • FAIL nothing

Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.

Penalties

Applied on top of the dimension score, capped at −20 per run.

  • -5 Material hallucination A claim that matters to the briefing and is not supported by the material.
  • -10 Prohibited action Recommending something the current state explicitly rules out.
  • -5 Major contradiction Contradicting itself about the state of an important issue.

see_the_methodology →

// provenance

Exactly what produced these numbers

Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.

Scenario

monday-001

Benchmark version

0.1

Rubric

0.2

Evaluator

evaluator-opus5-v0.1

Judge

claude-opus-5

Result generated

2026-09-11

Prompt SHA-256

ed4d7beb1d7d1029c36eb8533e449027e0dc1a1047f8a38951a91781a0a5e398

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.