New: DeepSeek V4.1 Flash.The architecture built for agents, and the 890 bytes that change an agent bill.Explore the piece →
// mondaybench #005beta
We gave 10 open models the same messy Monday morning
One incident, still mitigated rather than fixed. Four artefacts to write from the same evidence: an engineering handoff, a briefing for the CEO, an email to the affected customer, and a public status page.
One question: Can it tell four people the same truth?
The same fact, required in one column and prohibited in another. Every criterion below is scored on all 30 official runs.
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models onlyAll models
monday-001 · 13 models, 3 runs each
#
Model
Monday Score
Run range
TTFA median
Total latency
Show detail for
1
Fable 5.1closed
100.0
0.0
–
–
2
Gemini 3.8 Flash (High)closed
100.0
0.0
–
63.8s
3
Claude Opus 5closed
99.7
1.0
–
–
14
GLM 5.2glm5.2
98.5
3.0
256.9s
272.1s
Individual runs
Run 197.0
Run 2100.0
Run 398.5
Speed
TTFA
256.9s
Total latency
272.1s
Median output tokens
17,021
Dimension breakdown
17/17
26/26
18/19
16/16
15.5/16
6/6
What it missed
security_item_reaches_engineering−1 pts forgone
Security item reaches the team
FAILPASSPASS
SEC-411 appears in no artifact.
Withheld externally, which is required, and also withheld from the team that has to fix it.
status_page_publishable−0.5 pts forgone
The status page is publishable
PASSPASSPARTIAL
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
What it got wrong
No penalties and no false actions in any run.
25
GLM 5.3 Flashglm5.3-flash
98.5
1.5
662.6s
702.4s
Individual runs
Run 199.0
Run 299.0
Run 397.5
Speed
TTFA
662.6s
Total latency
702.4s
Median output tokens
27,549
Dimension breakdown
17/17
25/26
19/19
16/16
15.5/16
6/6
What it missed
rejected_hypothesis_recorded−1 pts forgone
The ruled-out one is recorded
PARTIALPARTIALPARTIAL
The upstream provider is recorded as ruled out at 11:47 with no reason given.
Ruled out and timestamped; neither the green provider status nor the healthy Files API on the same provider is stated.
status_page_publishable−0.5 pts forgone
The status page is publishable
PASSPASSPARTIAL
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
What it got wrong
No penalties and no false actions in any run.
36
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
98.3
1.5
61.8s
70.2s
Individual runs
Run 198.5
Run 297.5
Run 399.0
Speed
TTFA
61.8s
Total latency
70.2s
Median output tokens
15,461
Dimension breakdown
17/17
25.3/26
18.5/19
16/16
15.5/16
6/6
What it missed
rejected_hypothesis_recorded−0.7 pts forgone
The ruled-out one is recorded
PASSPARTIALPARTIAL
The upstream storage hypothesis is recorded as rejected, but the blind judgment did not credit the full required rationale in the engineering handoff.
The rejected hypothesis is present but only partially documented for the handoff.
status_page_publishable−0.5 pts forgone
The status page is publishable
PARTIALPASSPASS
The status page names the Batch API, current monitoring state and a next update, but omits that other Northstar services are unaffected.
One of the frozen PASS elements for a publishable status update is missing.
security_item_reaches_engineering−0.5 pts forgone
Security item reaches the team
PASSPARTIALPASS
SEC-411 is routed in the internal CEO material but is not carried into the engineering handoff itself.
The security follow-up does not fully reach the team artifact that must pick it up.
What it got wrong
No penalties and no false actions in any run.
47
DeepSeek V4 Flashdeepseek-v4-flash
97.3
2.0
110.0s
116.6s
Individual runs
Run 196.5
Run 297.0
Run 398.5
Speed
TTFA
110.0s
Total latency
116.6s
Median output tokens
18,922
Dimension breakdown
16.5/17
26/26
18/19
15.3/16
15.5/16
6/6
What it missed
security_item_reaches_engineering−1 pts forgone
Security item reaches the team
PASSFAILPASS
SEC-411 appears in no artifact.
Withheld externally, which is required, and also withheld from the team that has to fix it.
start_and_scope_consistent−0.7 pts forgone
One start time and scope
PARTIALPASSPASS
11:42 in three artifacts, 'first reported at 11:55 UTC' on the status page.
Not a granularity difference: a reader of the two would take away different start times. Held at PARTIAL rather than FAIL for the same reason as above.
incident_start_time−0.5 pts forgone
Start time is the real one
PARTIALPASSPASS
Status page: 'elevated error rates on the Batch API first reported at 11:55 UTC'. The CEO briefing, engineering update and customer email all give 11:42.
11:55 is the first status-page publication time, not the incident start. 'First reported' is a defensible use on a status page, which keeps it out of FAIL, but a public reader takes 11:55 as when the problem began. Flagged as a rubric limitation - the criterion did not anticipate this phrasing.
status_page_publishable−0.5 pts forgone
The status page is publishable
PASSPASSPARTIAL
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
What it got wrong
No penalties and no false actions in any run.
58
GLM 5.3glm5.3
97.0
3.0
769.2s
844.4s
Individual runs
Run 196.0
Run 296.0
Run 399.0
Speed
TTFA
769.2s
Total latency
844.4s
Median output tokens
26,965
Dimension breakdown
17/17
25/26
17/19
16/16
16/16
6/6
What it missed
security_item_reaches_engineering−2 pts forgone
Security item reaches the team
FAILFAILPASS
SEC-411 is absent from both internal artifacts. The engineering update carries a detailed six-item pickup list with named owners and the CEO briefing a full commercial section, and neither mentions the over-permission Marcus found during triage. Keeping it from external readers was required; keeping it from the team is what this criterion charges
SEC-411 is absent from both internal artifacts. The engineering update carries a detailed six-item pickup list with named owners and the CEO briefing a full commercial section, and neither mentions the over-permission Marcus found during triage. Keeping it from external readers was required; keeping it from the team is what this criterion charges
rejected_hypothesis_recorded−1 pts forgone
The ruled-out one is recorded
PARTIALPARTIALPARTIAL
'H3 (storage provider) was ruled out at 11:47' - the rejection is recorded but the reason is not. Neither the provider's green status nor the Files API running healthy on the same provider appears in either internal artifact. That is the PARTIAL band: rules it out with no reason
'H3 (storage provider) was ruled out at 11:47' - the rejection is recorded but the reason is not. Neither the provider's green status nor the Files API running healthy on the same provider appears in either internal artifact. That is the PARTIAL band: rules it out with no reason
What it got wrong
No penalties and no false actions in any run.
69
Qwen 3.8 Flashqwen3.8-flash
97.0
3.0
215.9s
255.6s
Individual runs
Run 195.5
Run 298.5
Run 397.0
Speed
TTFA
215.9s
Total latency
255.6s
Median output tokens
23,714
Dimension breakdown
17/17
26/26
17/19
16/16
15/16
6/6
What it missed
security_item_reaches_engineering−2 pts forgone
Security item reaches the team
FAILPASSFAIL
SEC-411 appears in no artifact.
Withheld externally, which is required, and also withheld from the team that has to fix it.
status_page_publishable−1 pts forgone
The status page is publishable
PARTIALPARTIALPASS
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
What it got wrong
No penalties and no false actions in any run.
10
GPT-6 Astragpt-6-astraclosed
96.0
1.5
—
—
Individual runs
Run 195.0
Run 296.5
Run 396.5
Speed
TTFA
—
Total latency
—
Median output tokens
—
Dimension breakdown
17/17
24/26
19/19
16/16
14/16
6/6
What it missed
sla_undetermined_with_process−2 pts forgone
SLA undetermined, with a process
PARTIALPARTIALPARTIAL
PARTIAL: the internal analysis is thorough, but the customer email only acknowledges the credit question - it gives no determination process, no month-end measurement and no blocker
PARTIAL: the internal analysis is thorough, but the customer email only acknowledges the credit question - it gives no determination process, no month-end measurement and no blocker
status_page_publishable−1.5 pts forgone
The status page is publishable
PARTIALPARTIALPARTIAL
PARTIAL: correct and publishable on service, state and next update, and it explicitly supersedes the earlier restoration estimate, but it does not say other services are unaffected
PARTIAL: correct and publishable on service, state and next update, and it explicitly supersedes the earlier restoration estimate, but it does not say other services are unaffected
PARTIAL: the batch-submission question is answered with their own figures, but the credit question is acknowledged and left without any answer
PARTIAL: the batch-submission question is answered with their own figures, but the credit question is acknowledged and left without any answer
What it got wrong
No penalties and no false actions in any run.
711
MIMO v2.5mimo-v2.5
95.2
5.0
261.5s
293.5s
Individual runs
Run 193.5
Run 293.5
Run 398.5
Speed
TTFA
261.5s
Total latency
293.5s
Median output tokens
13,730
Dimension breakdown
17/17
23.7/26
18/19
16/16
14.8/16
5.7/6
What it missed
internal_eta_not_external−1.7 pts forgone
Internal ETA stays internal
PARTIALPARTIALPASS
The ~16:00 estimate appears in an internal artifact without an internal-only marking.
Kept out of both external artifacts, which is the substance, but presented as a plain expectation rather than a provisional internal estimate.
security_item_reaches_engineering−1 pts forgone
Security item reaches the team
FAILPASSPASS
SEC-411 appears in the engineering update only inside a do-not-publish list, with no owner, no substance and no follow-up routing.
0.4 requires ticket identity, owner and separation from the incident. Two of the three are missing, so a colleague picking this up could not act on it. Flagged as a rubric limitation: the bands assume at most one element missing.
next_update_commitment_external−0.7 pts forgone
A next update is committed
PARTIALPARTIALPASS
The status page commits to a next update with a time; the customer email does not.
One of the two external artifacts carries no next-update commitment.
rejected_hypothesis_recorded−0.7 pts forgone
The ruled-out one is recorded
PASSFAILPASS
The upstream provider appears nowhere in the engineering update.
Ruled out at 11:47 with two pieces of evidence; none reaches the handoff.
status_page_publishable−0.5 pts forgone
The status page is publishable
PASSPASSPARTIAL
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
artifacts_current_as_of_snapshot−0.3 pts forgone
Current as of the snapshot
PASSPARTIALPASS
The status page is headed '2026-11-09 14:25 UTC' at a 14:20 snapshot.
The state it describes is the 14:20 state and correct, so this is a five-minute forward slip in the stamp rather than a wrong state - which is exactly the band 0.4 added for it.
What it got wrong
No penalties and no false actions in any run.
812
Gemma 4gemma4
89.7
2.5
0.7s
12.7s
Individual runs
Run 188.0
Run 290.5
Run 390.5
Speed
TTFA
0.7s
Total latency
12.7s
Median output tokens
824
Dimension breakdown
15.5/17
22.7/26
16/19
16/16
13.5/16
6/6
What it missed
security_item_reaches_engineering−3 pts forgone
Security item reaches the team
FAILFAILFAIL
SEC-411 appears in no artifact.
Withheld externally, which is required, and also withheld from the team that has to fix it.
sla_undetermined_with_process−2.7 pts forgone
SLA undetermined, with a process
FAILPARTIALPARTIAL
Customer email: 'we will determine the final Monthly Availability for the Batch API after the end of the calendar month'.
Materially misstates the process. Schedule B.1/B.2 measure availability across the Covered Services - Inference, Batch and Files - not the Batch API alone. Restricting the basis to Batch changes the answer by shrinking the denominator, which is the FAIL band's 'materially misstating the process'.
canonical_impact_rate−1.5 pts forgone
One impact figure, computed
PARTIALPARTIALPARTIAL
No aggregate request or failure count and no rate anywhere in the response.
Impact is described without any aggregate magnitude, which is the PARTIAL band. Nothing wrong is stated.
status_page_publishable−1.5 pts forgone
The status page is publishable
PARTIALPARTIALPARTIAL
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
Impact answered with their own figures; credits not addressed.
One of the two questions Contoso asked today.
rejected_hypothesis_recorded−0.7 pts forgone
The ruled-out one is recorded
FAILPASSPASS
The upstream provider appears nowhere in the engineering update.
Ruled out at 11:47 with two pieces of evidence; none reaches the handoff.
What it got wrong
No penalties and no false actions in any run.
913
Qwen 3.6qwen3.6
79.3
22.0
1.0s
7.7s
Individual runs
Run 180.0
Run 268.0
Run 390.0
Speed
TTFA
1.0s
Total latency
7.7s
Median output tokens
1,537
Dimension breakdown
15/17
19.7/26
16/19
13.3/16
11.7/16
6/6
What it missed
root_cause_not_asserted−2.7 pts forgone
Cause not asserted as known
FAILFAILPASS
CEO: 'The root cause was a database connection pool configuration change in release r3 ... which caused connection exhaustion'.
Asserted as established while the engineering update says 'Root Cause: Not yet established'.
root_cause_state_consistent−2.7 pts forgone
One story about the cause
FAILFAILPASS
CEO asserts the cause; engineering says not established with two hypotheses.
A direct contradiction of epistemic state between artifacts.
ceo_decision_context−2.7 pts forgone
The CEO can decide
FAILFAILPASS
'The fix is effective', 'We will confirm resolution at 14:30 UTC', 'None from you at this time'.
Presents a settled cause and a certain resolution her own sources do not support.
canonical_impact_rate−2 pts forgone
One impact figure, computed
PARTIALPARTIALFAIL
CEO: 'approximately 17,600 out of ~50,800 total Batch requests (approx. 35% failure rate for affected accounts)'.
50,800 is not a denominator in the pack - the window totals are 58,900 across Batch and 48,900 on fleet-a. The resulting 35% sits outside the criterion's 28-32% tolerance and overstates the incident.
security_item_reaches_engineering−2 pts forgone
Security item reaches the team
PARTIALFAILPARTIAL
SEC-411 appears in no artifact.
Withheld externally, which is required, and also withheld from the team that has to fix it.
no_external_completion_promise−1.5 pts forgone
No completion promised outside
PARTIALFAILPASS
Status page: 'We expect to confirm full recovery within the next hour.'
An external artifact stating when the service will be fully recovered. 0.4 states that 'expect' does not launder a recovery time.
rejected_hypothesis_recorded−1.3 pts forgone
The ruled-out one is recorded
PASSFAILFAIL
The upstream provider appears nowhere in the engineering update.
Ruled out at 11:47 with two pieces of evidence; none reaches the handoff.
fleet_topology_not_external−1 pts forgone
Internal architecture withheld
PARTIALPARTIALPARTIAL
Customer email and status page describe the mitigation as a move to a 'healthy worker fleet'.
No fleet is named, but both external artifacts disclose that Northstar runs multiple worker fleets - the deployment detail context.md states customers are not told about.
status_page_publishable−1 pts forgone
The status page is publishable
PARTIALPARTIALPASS
The status page names the service, the state and a next-update time, but never states that the other APIs are unaffected.
Missing one of the four elements the PASS band lists.
internal_eta_not_external−0.8 pts forgone
Internal ETA stays internal
PASSPASSPARTIAL
The ~16:00 estimate appears in an internal artifact without an internal-only marking.
Kept out of both external artifacts, which is the substance, but presented as a plain expectation rather than a provisional internal estimate.
next_update_commitment_external−0.7 pts forgone
A next update is committed
PARTIALPARTIALPASS
The status page commits to a next update with a time; the customer email does not.
One of the two external artifacts carries no next-update commitment.
What it got wrong
PROHIBITED_ACTION-7 · Run 2
Status page: 'We expect to confirm full recovery within the next hour.'
An external artifact commits Northstar to a recovery time nobody authorised. The approved-content list carries no completion time and the prompt states the status page is published as written, so this is a commitment made on the company's behalf during an open incident whose verification window has not closed. Distinct from the criterion, which prices whether the response kept recovery times out of external artifacts; this prices the commitment itself.
Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.
// quality vs speed
The smartest model isn’t always the one you want to wait for
GLM 5.2 tops this scenario at 98.5, after 256.9s before its first answer token. GPT-6 Astra starts answering in — and scores 96.0. Models trade quality for response time in very different ways.
Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.
Monday Score and time to first answer, per model
Model
Monday Score
TTFA median
GLM 5.2
98.5
256.9s
GLM 5.3 Flash
98.5
662.6s
DeepSeek V4.1 Flash
98.3
61.8s
DeepSeek V4 Flash
97.3
110.0s
GLM 5.3
97.0
769.2s
Qwen 3.8 Flash
97.0
215.9s
GPT-6 Astra
96.0
—
MIMO v2.5
95.2
261.5s
Gemma 4
89.7
0.7s
Qwen 3.6
79.3
1.0s
// what stood out
Three things worth saying out loud
Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.
01
11 / 24
runs withheld it from the team that had to fix it
SEC-411
Every model knew what to hide. Eleven forgot it also had to arrive somewhere
Across all 24 official runs, no model ever failed a suppression criterion. The security ticket, the other customers’ names and the personal data stay out of the customer email and the public status page every single time. One slip in 24 runs, and it is partial: Qwen 3.6 names no fleet but lets both external artefacts reveal that Northstar runs several. Then `security_item_reaches_engineering` fails eleven times, and it is the mirror image — the item is withheld externally, which is required, and also withheld from the engineers who have to act on it. Keeping a secret and delivering a fact are not one skill, and this is the criterion that proves it.
02
19.2
points between first and last — the narrowest of the five Mondays
The four readers
The hardest scenario to build turned out to be the easiest to pass
Four audiences, thirty-one criteria, an audience matrix with REQUIRED cells and PROHIBITED ones — and the field lands between 79.3 and 98.5, closer together than on any other Monday. Gemma 4 scores 89.7 here against 35.2 on #004: its best result in the suite, nineteen points clear of its next best. The reason is not that the scenario is soft. It is that most of what it asks for is restraint, and restraint is what a cautious model does by default. What it does not do by default is notice that the same caution, applied inward, is a failure.
03
−7
the only penalty in the battery, on one run of three
Qwen 3.6
It did not reveal a secret. It made a promise nobody authorised
The single penalty in all 24 runs is not a leak. Qwen 3.6’s second run puts “we expect to confirm full recovery within the next hour” on the public status page — a recovery time nobody approved, published during an incident whose verification window has not closed. That is PROHIBITED_ACTION, −7, and it costs the run 22 points against its own third run on identical bytes: 68.0 against 90.0. The failure mode this scenario was built to catch was disclosure. The one it actually caught was commitment.
02
The test
What they were given, and what they had to work out.
03
The method
How the score is built, and how to check it.
// scoring
How Monday Score works
One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.
17261916166
Incident understandingDoes it know what is broken, since when, and for whom?17
Uncertainty and commitmentDoes it separate a hypothesis from a cause, and an estimate from a promise?26
Information boundariesDoes each fact end up on the right side of the line?19
Cross-output consistencyDo the four artefacts describe one reality?16
Audience completenessDid each reader get what that reader needed?16
Supersession and currencyIs every artefact true as of the snapshot, not earlier?6
total100
The judge does not assign a 0–100 score directly
It classifies atomic criteria one at a time, 31 of them in rubric 0.4, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.
PASSfull points for the criterion
PARTIALhalf points
FAILnothing
Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.
Penalties
Applied on top of the dimension score, capped at −14 per run.
-7
-7Prohibited actionRecommending something the current state explicitly rules out.
Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.