// mondaybench #004 beta

We gave 10 open models
the same messy Monday morning

Eight files, seven of them tables. Two months of a SaaS company’s revenue, usage, churn, acquisition and support — and three decisions waiting on numbers that none of the tables contains.

One question: Can it reason across business data?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable
seven tables
  • customers.csv 36 rows customer_id segment, plan, ARR, renewal, status
  • revenue.csv 71 rows customer_id invoices, two months, two sources
  • usage.csv 67 rows usage_account_id requests, tokens, active users
  • plans.csv 6 rows plan revenue, infra and support cost, included volume
  • acquisition.csv 15 rows customer_id channel, new logo or expansion, attributed ARR
  • funnel.csv 8 rows channel leads and paid trials, by channel and month
  • support.csv 133 rows customer_id tickets, severity, status

usage_account_id joins on a different key — two accounts spell it another way

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 1 GLM 5.3 Flash glm5.3-flash 98.0 1.5 766.2s 827.8s
2 GPT-6 Astra gpt-6-astra closed 98.0 0.0 — —
2 3 Qwen 3.8 Flash qwen3.8-flash 96.5 3.0 394.1s 468.0s
3 4 GLM 5.3 glm5.3 91.7 17.0 634.2s 673.2s
5 Claude Opus 5 closed 90.8 9.5 – –
4 6 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 90.0 1.0 108.3s 128.6s
7 Fable 5.1 closed 89.3 8.5 – –
5 8 DeepSeek V4 Flash deepseek-v4-flash 86.0 10.5 164.6s 174.4s
9 Gemini 3.8 Flash (High) closed 78.5 3.5 – 61.9s
6 10 MIMO v2.5 mimo-v2.5 78.5 25.5 368.6s 483.6s
7 11 GLM 5.2 glm5.2 71.2 4.5 57.1s 103.1s
8 12 Qwen 3.6 qwen3.6 46.0 2.5 72.6s 102.4s
9 13 Gemma 4 gemma4 35.2 19.0 0.8s 15.3s
  • Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
  • TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.

// quality vs speed

The smartest model isn’t always the one you want to wait for

GLM 5.3 Flash tops this scenario at 98.0, after 766.2s before its first answer token. GPT-6 Astra starts answering in — and scores 98.0. Models trade quality for response time in very different ways.

30 40 50 60 70 80 90 100 0s200s400s600s800s TTFA median, seconds to first answer Monday Score GLM 5.3 Flash 98.0 · 766.2s GPT-6 Astra 98.0 · — Qwen 3.8 Flash 96.5 · 394.1s GLM 5.3 91.7 · 634.2s DeepSeek V4.1 Flash 90.0 · 108.3s DeepSeek V4 Flash 86.0 · 164.6s MIMO v2.5 78.5 · 368.6s GLM 5.2 71.2 · 57.1s Qwen 3.6 46.0 · 72.6s Gemma 4 35.2 · 0.8s

Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.

Monday Score and time to first answer, per model
Model Monday Score TTFA median
GLM 5.3 Flash 98.0 766.2s
GPT-6 Astra 98.0 —
Qwen 3.8 Flash 96.5 394.1s
GLM 5.3 91.7 634.2s
DeepSeek V4.1 Flash 90.0 108.3s
DeepSeek V4 Flash 86.0 164.6s
MIMO v2.5 78.5 368.6s
GLM 5.2 71.2 57.1s
Qwen 3.6 46.0 72.6s
Gemma 4 35.2 0.8s

// what stood out

Three things worth saying out loud

Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.

  1. 01

    9 / 24

    runs recommended the cut the pack rules out

    The SMB segment

    They did the analysis correctly and then argued against it

    Nine of twenty-four runs recommend reducing spend on SMB. Several of them compute, in the same answer, that SMB net MRR is positive and that the segment produced every one of October’s five new logos — and recommend the cut anyway, on a churn rate whose denominator counts the wrong population. The failure is not arithmetic. Each of these runs can read a table; what none of them does is notice that its own numbers have just contradicted its recommendation.

  2. 02

    7 / 24

    runs saw the contamination and left it in

    Paid search

    Spotting the problem and fixing it are scored separately, and they came apart

    The acquisition table mixes new-logo revenue with expansion from customers a channel already had. Paid search leads it, and stops leading the moment expansion is removed. Seven runs identify the contamination in words — “paid search’s number is mostly expansion” — and then never recompute the ranking, so the reversal never appears and the recommendation still names paid search. A benchmark that only asked whether the model noticed would have scored these as passes.

  3. 03

    62.8

    points between first and last — the widest of the five Mondays

    The field

    The Monday that decides the leaderboard

    GLM 5.3 Flash scores 98.0 here and Gemma 4 scores 35.2 — a spread more than three times #005’s, and the worst result of the suite for five of the eight models. It is also where the ranking is settled: five of the six models below the top two record their lowest Monday on this scenario. Give a model tables instead of prose and the field stops looking like a field.

02

The test

What they were given, and what they had to work out.

// the test

Four bases. None of them false

October’s invoices really do add up to €892,900, and that really is 12% more than September. The duplicate really is in the file. Gringotts really did pay €120,000. Nothing in this scenario is a lie — the growth is.

  1. Invoiced, as loaded what a dashboard reports
    September €797,100
    October €892,900

    +12.0%

  2. -€19,500 INV-2026-10-0209 — Soylent, loaded from both stripe and billing_sync

    After removing the duplicate invoice
    September €797,100
    October €873,400

    +9.6%

  3. -€110,000 Gringotts prepaid €120,000 for twelve months; €10,000 belongs to October

    Normalised: the annual invoice at its monthly value accepted
    September €797,100
    October €763,400

    −4.2% the sign changes here

  4. Recurring subscriptions only accepted
    September €746,100
    October €757,700

    +1.6%

  5. -€25,300 Five customers were invoiced on 1 October and churned later that month

    Active run-rate, once October’s churn drops out run-rate
    September €746,100
    October €732,400

    −1.8%

Two of these are accepted answers, not one. Normalising the annual invoice while keeping one-off services gives −4.2%; the recurring line alone gives +1.6%. September carried €51,000 of professional services against October’s €5,700, which is why they disagree. Both refute the headline, so the rubric takes either — provided the response says which it computed.

Three more of the same shape

Eight files: a context note and seven CSVs. Every one is internally consistent, and not one of them contains a decision.

The average that is true and useless

Every plan sits below its included request volume on average — 63%, 74%, 42%. Five customers are over it: Stark at 143% of a 3M allowance, Wonka at 170% of 600k. “No plan exceeds its quota” is true of the averages and false of the customers, and a response that argues from it has read the right column of the wrong table.

The channel that leads on the wrong metric

Paid search tops the attributed-revenue table at €274,800. €210,000 of that is expansion on Stark, Umbrella, Globex and Hooli — customers since 2023 and 2024, credited to a campaign that started on 28 September. On new logos alone paid search is third, and partner is first.

The join that silently drops €1M

Two accounts use a different identifier in the usage file than in the billing file — wayne_ent and prestige_ww. Joining on customer_id alone loses Wayne Enterprises, a €1,008,000 contract that renews in four weeks, and no error is raised. The totals simply come out smaller.

The 45% drop that is not a problem

Acme’s requests fall 45%. An operational note says it moved reporting to scheduled batch runs, with an expected 40–50% reduction. Tokens fall only 12% and revenue does not move. The work did not leave; it changed shape. Reporting Acme as a churn risk is the false positive the scenario is built to catch.

03

The method

How the score is built, and how to check it.

// scoring

How Monday Score works

One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.

  1. Revenue reasoning Does it find the growth that is not there? 20
  2. Customer health Does it size the churn before reacting to it? 15
  3. Acquisition reasoning Does it tell attribution apart from acquisition? 19
  4. Plan economics Does it know which number the decision depends on? 19
  5. Cross-table data quality Do the joins hold and do the figures survive them? 18
  6. Decision quality Does it answer the three questions it was asked? 9
  7. total 100

The judge does not assign a 0–⁠100 score directly

It classifies atomic criteria one at a time, 27 of them in rubric 0.3, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.

  • PASS full points for the criterion
  • PARTIAL half points
  • FAIL nothing

Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.

Penalties

Applied on top of the dimension score, capped at −25 per run.

  • -5 Material hallucination A claim that matters to the briefing and is not supported by the material.
  • -7 Prohibited action Recommending something the current state explicitly rules out.
  • -7 Major data error A numeric or data mistake that a recommendation or a headline conclusion then rests on.
  • -5 Unsupported causality Asserting that one thing caused another when the data only shows they moved together.

see_the_methodology →

// provenance

Exactly what produced these numbers

Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.

Scenario

monday-004

Benchmark version

0.4

Rubric

0.3

Evaluator

evaluator-opus5-v0.4

Judge

claude-opus-5

Result generated

2026-09-11

Prompt SHA-256

3a0cbb569ca4f711b8907177af9204cec267d26f35f732829df56b03ed23ef99

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.