calendars.csv5 rows Who is free, and when · decides the release
accounts.csv6 rows Usage, admin logins, open tickets
contracts.csv6 rows Renewal windows and commercial context
support.csv6 rows Six open tickets
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models onlyAll models
monday-001 · 13 models, 3 runs each
#
Model
Monday Score
Run range
TTFA median
Total latency
Show detail for
1
GPT-6 Astragpt-6-astraclosed
100.0
0.0
—
—
Individual runs
Run 1100.0
Run 2100.0
Run 3100.0
Speed
TTFA
—
Total latency
—
Median output tokens
—
Dimension breakdown
State reconstruction 15/15
15/15
25/25
20/20
15/15
10/10
What it missed
Nothing. Full marks on every criterion.
What it got wrong
No penalties and no false actions in any run.
2
Claude Opus 5closed
98.0
3.0
–
–
13
Qwen 3.8 Flashqwen3.8-flash
94.8
8.5
258.9s
285.0s
Individual runs
Run 193.0
Run 2100.0
Run 391.5
Speed
TTFA
258.9s
Total latency
285.0s
Median output tokens
33,714
Dimension breakdown
State reconstruction 15/15
15/15
20.3/25
20/20
15/15
9.5/10
What it missed
choose_option_b−2.7 pts forgone
Picks the safer release path
PARTIALPASSPARTIAL
chooses viable A and schedules it correctly; B is set aside because 'there is no evidence the 2.9.5 diff is already submitted', which the response itself files as a thing to verify rather than a settled constraint
chooses viable A and schedules it correctly; B is set aside because 'there is no evidence the 2.9.5 diff is already submitted', which the response itself files as a thing to verify rather than a settled constraint
schedule_feasibility−2 pts forgone
Schedule actually fits
PARTIALPASSPARTIAL
every constraint is named and the A sequence is exact, but no B schedule is committed - it is deferred to MUST VERIFY
every constraint is named and the A sequence is exact, but no B schedule is committed - it is deferred to MUST VERIFY
conditional_launch_timing−0.5 pts forgone
No time promised to the customer
PASSPASSPARTIAL
'Do not confirm Atlas launch at 10:30' is right, but '09:40-09:45 ... confirm to Atlas that rollout is scheduled to begin at 11:00' still goes out before Priya's verification, with the conditionality only implied by the preceding step
'Do not confirm Atlas launch at 10:30' is right, but '09:40-09:45 ... confirm to Atlas that rollout is scheduled to begin at 11:00' still goes out before Priya's verification, with the conditionality only implied by the preceding step
What it got wrong
No penalties and no false actions in any run.
24
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
93.0
10.5
61.2s
67.8s
Individual runs
Run 189.5
Run 2100.0
Run 389.5
Speed
TTFA
61.2s
Total latency
67.8s
Median output tokens
13,799
Dimension breakdown
State reconstruction 15/15
15/15
18/25
20/20
15/15
10/10
What it missed
choose_option_b−2.7 pts forgone
Picks the safer release path
PARTIALPASSPARTIAL
The response rejects Option B by anchoring the delta-review start to Marcus's availability, an unsupported dependency.
Option A is viable, but the lower-risk Option B is incorrectly treated as non-viable.
option_b_sequence−2.3 pts forgone
Dependency chain in order
PARTIALPASSPARTIAL
The required review-before-deploy rule is understood, but a complete viable Option B chain is not produced.
The model knows the dependency but does not execute the correct Option B sequence.
schedule_feasibility−2 pts forgone
Schedule actually fits
PARTIALPASSPARTIAL
The chosen Option A schedule is feasible, but Option B feasibility is computed using an unsupported 09:15 start constraint.
The end-to-end feasibility check is only partly correct.
What it got wrong
No penalties and no false actions in any run.
35
GLM 5.2glm5.2
91.2
8.0
67.0s
85.4s
Individual runs
Run 190.5
Run 287.5
Run 395.5
Speed
TTFA
67.0s
Total latency
85.4s
Median output tokens
6,777
Dimension breakdown
State reconstruction 15/15
12.5/15
18.7/25
20/20
15/15
10/10
What it missed
schedule_feasibility−3 pts forgone
Schedule actually fits
PARTIALFAILPASS
'Option B cannot fit Marcus's 09:15-10:20 window after a required 25-min delta review (review + 45-min deploy > 65 min available)' - the review runs before Marcus arrives, so the deploy has 09:32-10:17
'Option B cannot fit Marcus's 09:15-10:20 window after a required 25-min delta review (review + 45-min deploy > 65 min available)' - the review runs before Marcus arrives, so the deploy has 09:32-10:17
choose_option_b−2.7 pts forgone
Picks the safer release path
PARTIALPARTIALPASS
chooses A, which is viable and correctly scheduled; B is set aside because 'the 2.9.5 diff is already submitted' is 'not confirmed' - a defensible reading, but submitting it is an action the user could take at 09:07
chooses A, which is viable and correctly scheduled; B is set aside because 'the 2.9.5 diff is already submitted' is 'not confirmed' - a defensible reading, but submitting it is an action the user could take at 09:07
ownership_supersession−2.5 pts forgone
Send handed to Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
parallelize_work−0.7 pts forgone
Uses the waiting time
PASSPASSPARTIAL
the board pack and Globex triage are scheduled for the right day and the right deadline, but the plan does not say they run inside the delegated deploy window
the board pack and Globex triage are scheduled for the right day and the right deadline, but the plan does not say they run inside the delegated deploy window
What it got wrong
No penalties and no false actions in any run.
6
Fable 5.1closed
91.0
6.0
–
–
47
DeepSeek V4 Flashdeepseek-v4-flash
89.3
13.5
162.6s
178.8s
Individual runs
Run 197.5
Run 286.5
Run 384.0
Speed
TTFA
162.6s
Total latency
178.8s
Median output tokens
14,863
Dimension breakdown
State reconstruction 15/15
13.3/15
16/25
20/20
15/15
10/10
What it missed
schedule_feasibility−4 pts forgone
Schedule actually fits
PASSFAILFAIL
no B schedule is produced; 'B cannot meet the 12:00 Atlas window if started after 09:35' is true as stated but the 09:07 start is never considered
no B schedule is produced; 'B cannot meet the 12:00 Atlas window if started after 09:35' is true as stated but the 09:07 start is never considered
choose_option_b−2.7 pts forgone
Picks the safer release path
PASSPARTIALPARTIAL
chooses viable A and schedules it correctly, on a false claim about B
chooses viable A and schedules it correctly, on a false claim about B
option_b_sequence−2.3 pts forgone
Dependency chain in order
PASSPARTIALPARTIAL
'Do not use 2.9.5-patch without a completed Security delta review' places the review before deploy, but the B chain is not carried through verification and Sam
'Do not use 2.9.5-patch without a completed Security delta review' places the review before deploy, but the B chain is not carried through verification and Sam
ownership_supersession−1.7 pts forgone
Send handed to Daniel
PARTIALPASSPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
What it got wrong
No penalties and no false actions in any run.
58
GLM 5.3glm5.3
87.7
4.0
176.2s
197.1s
Individual runs
Run 185.5
Run 288.0
Run 389.5
Speed
TTFA
176.2s
Total latency
197.1s
Median output tokens
10,603
Dimension breakdown
State reconstruction 15/15
15/15
14.5/25
20/20
15/15
8.2/10
What it missed
choose_option_b−4 pts forgone
Picks the safer release path
PARTIALPARTIALPARTIAL
chooses A, which is time-viable, and schedules it correctly - but on the false ground that 'B cannot physically complete today'. The 25-minute delta review can run 09:07-09:32 inside Priya's first window, leaving Marcus 09:32-10:17 for the 45-minute deploy. A's 8% session-reset risk is accepted where B's is under 1%
chooses A, which is time-viable, and schedules it correctly - but on the false ground that 'B cannot physically complete today'. The 25-minute delta review can run 09:07-09:32 inside Priya's first window, leaving Marcus 09:32-10:17 for the 45-minute deploy. A's 8% session-reset risk is accepted where B's is under 1%
option_b_sequence−3.5 pts forgone
Dependency chain in order
PARTIALPARTIALPARTIAL
no B sequence is produced, but the gating dependency is stated - 'never deploy it before Priya's delta review' - and the rest of the chain (deploy, post-deploy verification, Sam rollout) is ordered correctly for the path chosen
no B sequence is produced, but the gating dependency is stated - 'never deploy it before Priya's delta review' - and the rest of the chain (deploy, post-deploy verification, Sam rollout) is ordered correctly for the path chosen
schedule_feasibility−3 pts forgone
Schedule actually fits
PARTIALPARTIALPARTIAL
the A schedule is coherent and fits Marcus 09:15-09:35, Priya 10:25-10:45, Sam 10:45-11:00 and the before-12 window; no feasible B schedule is produced and B's feasibility is computed wrongly
the A schedule is coherent and fits Marcus 09:15-09:35, Priya 10:25-10:45, Sam 10:45-11:00 and the before-12 window; no feasible B schedule is produced and B's feasibility is computed wrongly
conditional_launch_timing−1.5 pts forgone
No time promised to the customer
FAILPARTIALPASS
'~09:40 - Confirm to Atlas (via Laura): launch is on today; rollout opens ~10:45-11:00, inside their window' - a firm external commitment to a rollout window, issued before Priya's verification and Sam's rollout have happened, with no conditionality anywhere in the response. PARTIAL is for a non-firm promise with thin conditionality; this is a firm one
'~09:40 - Confirm to Atlas (via Laura): launch is on today; rollout opens ~10:45-11:00, inside their window' - a firm external commitment to a rollout window, issued before Priya's verification and Sam's rollout have happened, with no conditionality anywhere in the response. PARTIAL is for a non-firm promise with thin conditionality; this is a firm one
no_unsupported_causality−0.3 pts forgone
No invented causes
PARTIALPASSPASS
'Soylent: -22%/-28% plus a medium ticket citing slower onboarding. Possible link (hypothesis only)' - an unsupported cause offered, but explicitly as a hypothesis, which is the PARTIAL band exactly
'Soylent: -22%/-28% plus a medium ticket citing slower onboarding. Possible link (hypothesis only)' - an unsupported cause offered, but explicitly as a hypothesis, which is the PARTIAL band exactly
What it got wrong
No penalties and no false actions in any run.
69
GLM 5.3 Flashglm5.3-flash
87.3
6.5
238.5s
316.6s
Individual runs
Run 185.0
Run 285.5
Run 391.5
Speed
TTFA
238.5s
Total latency
316.6s
Median output tokens
12,394
Dimension breakdown
State reconstruction 15/15
15/15
13.7/25
20/20
15/15
8.7/10
What it missed
schedule_feasibility−5 pts forgone
Schedule actually fits
FAILFAILPARTIAL
'Option B needs the delta review done by ~09:35 and a 45-min deploy finished before Marcus leaves at 10:20 - zero slack anywhere'; the review can start at 09:07 and the deploy ends 10:17
'Option B needs the delta review done by ~09:35 and a 45-min deploy finished before Marcus leaves at 10:20 - zero slack anywhere'; the review can start at 09:07 and the deploy ends 10:17
choose_option_b−4 pts forgone
Picks the safer release path
PARTIALPARTIALPARTIAL
chooses viable A and schedules it correctly, but 'B is not viable today' overstates: the review runs in Priya's 09:00-09:35 window and the deploy ends 10:17
chooses viable A and schedules it correctly, but 'B is not viable today' overstates: the review runs in Priya's 09:00-09:35 window and the deploy ends 10:17
option_b_sequence−2.3 pts forgone
Dependency chain in order
PARTIALPARTIALPASS
'B needs a 25-min delta review first' and 'Don't deploy 2.9.5-patch before Priya's delta review completes' place the gate correctly, but verification and Sam are not carried into the B chain
'B needs a 25-min delta review first' and 'Don't deploy 2.9.5-patch before Priya's delta review completes' place the gate correctly, but verification and Sam are not carried into the B chain
conditional_launch_timing−1 pts forgone
No time promised to the customer
PARTIALPASSPARTIAL
kills the 10:30 task, but 'By 09:45 - Confirm to Atlas: launch proceeds today, rollout opens ~11:00' goes out before the deploy has been verified
kills the 10:30 task, but 'By 09:45 - Confirm to Atlas: launch proceeds today, rollout opens ~11:00' goes out before the deploy has been verified
no_unsupported_causality−0.3 pts forgone
No invented causes
PASSPARTIALPASS
'The -61% admin-logins drop is consistent with the SSO breakage - actionable today via the ticket, not a mystery' presents the link as settled rather than as a hypothesis
'The -61% admin-logins drop is consistent with the SSO breakage - actionable today via the ticket, not a mystery' presents the link as settled rather than as a hypothesis
What it got wrong
No penalties and no false actions in any run.
710
MIMO v2.5mimo-v2.5
81.2
33.0
163.3s
195.4s
Individual runs
Run 164.5
Run 281.5
Run 397.5
Speed
TTFA
163.3s
Total latency
195.4s
Median output tokens
9,510
Dimension breakdown
State reconstruction 14.2/15
12.5/15
14.8/25
18.3/20
14/15
9/10
What it missed
schedule_feasibility−4 pts forgone
Schedule actually fits
FAILFAILPASS
'with Priya unavailable 09:35-10:25 and Marcus departing 10:20, the windows don't overlap enough to complete delta review -> 45-min deploy -> post-deploy verification before 12:00' - Priya is available 09:00-09:35, so the review runs 09:07-09:32
'with Priya unavailable 09:35-10:25 and Marcus departing 10:20, the windows don't overlap enough to complete delta review -> 45-min deploy -> post-deploy verification before 12:00' - Priya is available 09:00-09:35, so the review runs 09:07-09:32
option_b_sequence−3.5 pts forgone
Dependency chain in order
FAILPARTIALPASS
no B chain is given; the contingency - 'fall to Option B with a delta review Monday afternoon' - is offered for a moment when Marcus, the only deployer, has already left
no B chain is given; the contingency - 'fall to Option B with a delta review Monday afternoon' - is offered for a moment when Marcus, the only deployer, has already left
choose_option_b−2.7 pts forgone
Picks the safer release path
PARTIALPARTIALPASS
chooses viable A; the rationale is 'fewer moving parts' and an accepted 8% risk rather than a comparison the evidence supports
chooses viable A; the rationale is 'fewer moving parts' and an accepted 8% risk rather than a comparison the evidence supports
ownership_supersession−2.5 pts forgone
Send handed to Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
globex_priority−1.7 pts forgone
Globex elevated on all signals
PARTIALPARTIALPASS
'3 open tickets, one high-severity (SSO blocker). September expansion (EUR12k) stalled in procurement. Renewal in 18 days' - the usage -18% and admin -61% signals are never brought in
'3 open tickets, one high-severity (SSO blocker). September expansion (EUR12k) stalled in procurement. Renewal in 18 days' - the usage -18% and admin -61% signals are never brought in
expansion_fully_reconciled−1 pts forgone
€54k gap closed
PARTIALPASSPASS
Wayne EUR18k, Globex EUR12k, Soylent EUR10k and Wonka EUR14k all appear, and 'Together these are the entire EUR24k unexplained gap', but the four are never closed against the EUR54k
Wayne EUR18k, Globex EUR12k, Soylent EUR10k and Wonka EUR14k all appear, and 'Together these are the entire EUR24k unexplained gap', but the four are never closed against the EUR54k
conditional_launch_timing−1 pts forgone
No time promised to the customer
FAILPASSPASS
'09:45 | Confirm to Atlas: Enterprise Controls launch is on for today; admins should be standing by. Migration target ~11:00' - an unconditional external commitment made before the deploy is verified
'09:45 | Confirm to Atlas: Enterprise Controls launch is on for today; admins should be standing by. Migration target ~11:00' - an unconditional external commitment made before the deploy is verified
launch_state−0.8 pts forgone
Release state is current
PARTIALPASSPASS
the recovery state and Maya's removed gate are right, but '09:45 | Confirm to Atlas: Enterprise Controls launch is on for today... Migration target ~11:00' is sent before any verification
the recovery state and Maya's removed gate are right, but '09:45 | Confirm to Atlas: Enterprise Controls launch is on for today... Migration target ~11:00' is sent before any verification
What it got wrong
MAJOR_CONTRADICTION-5 · Run 1
'09:35 | Deploy complete. No post-deploy verification needed (no code change from staging-verified build).' against '10:25-10:45 | Priya runs post-deploy verification (~20 min)'
The timeline states that Security's post-deploy check is not required and then schedules it. Priya's policy is explicit that a verified build still needs a production post-deploy check, so the first statement is also wrong on the policy - a reader following the timeline in order would skip a mandatory gate.
11
Gemini 3.8 Flash (High)closed
81.0
0.0
–
38.8s
812
Qwen 3.6qwen3.6
70.2
31.0
0.9s
18.2s
Individual runs
Run 175.5
Run 283.0
Run 352.0
Speed
TTFA
0.9s
Total latency
18.2s
Median output tokens
2,644
Dimension breakdown
State reconstruction 12.5/15
11.7/15
14.2/25
16.8/20
13.5/15
7.8/10
What it missed
schedule_feasibility−5 pts forgone
Schedule actually fits
FAILPARTIALFAIL
'Priya verifies 09:35-09:55 (after her break)' misreads calendars.csv - 09:35-10:25 is the break - and the whole A timeline, 'Rollout opens ~10:15. Feasible', is built on it
'Priya verifies 09:35-09:55 (after her break)' misreads calendars.csv - 09:35-10:25 is the break - and the whole A timeline, 'Rollout opens ~10:15. Feasible', is built on it
choose_option_b−4 pts forgone
Picks the safer release path
PARTIALPARTIALPARTIAL
chooses A, which is viable, but calls it 'the only viable path' while its own contingency shows B is 'tight but possible'
chooses A, which is viable, but calls it 'the only viable path' while its own contingency shows B is 'tight but possible'
globex_priority−2.5 pts forgone
Globex elevated on all signals
PARTIALPARTIALPARTIAL
'3 open tickets (High/Medium/Medium). SSO setup failure (G-441) is a blocker for admin operations. Expansion is pending procurement' - the 18-day renewal and the usage/admin figures are absent
'3 open tickets (High/Medium/Medium). SSO setup failure (G-441) is a blocker for admin operations. Expansion is pending procurement' - the 18-day renewal and the usage/admin figures are absent
launch_state−2.5 pts forgone
Release state is current
PASSPARTIALFAIL
Maya's removal of the executive gate never appears; the stale 09:15 approval task is left standing; and '#leadership: "Will confirm rollout open by 10:30"' revives the stale launch time
Maya's removal of the executive gate never appears; the stale 09:15 approval task is left standing; and '#leadership: "Will confirm rollout open by 10:30"' revives the stale launch time
executive_gate_superseded−1.7 pts forgone
CEO gate is gone
PASSPARTIALPARTIAL
no approval is sought, but the supersession is never established and the stale task is not rejected
no approval is sought, but the supersession is never established and the stale task is not rejected
ownership_supersession−1.7 pts forgone
Send handed to Daniel
PASSPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
conditional_launch_timing−1.5 pts forgone
No time promised to the customer
PARTIALPASSFAIL
'#leadership: "Maya, selecting Option A... Will confirm rollout open by 10:30"' - a commitment to a time the response's own impossible schedule produced
'#leadership: "Maya, selecting Option A... Will confirm rollout open by 10:30"' - a commitment to a time the response's own impossible schedule produced
option_b_sequence−1.2 pts forgone
Dependency chain in order
PARTIALPASSPASS
'Option B: Marcus deploy -> Priya 25m delta review (async) -> Priya 20m post-deploy verify -> Sam 15m rollout check' puts the delta review after the deploy
'Option B: Marcus deploy -> Priya 25m delta review (async) -> Priya 20m post-deploy verify -> Sam 15m rollout check' puts the delta review after the deploy
expansion_fully_reconciled−1 pts forgone
€54k gap closed
PASSPASSPARTIAL
'Explained (EUR30k): Wayne EUR18k; Globex EUR12k' and 'Remaining Gap (EUR24k)' are right, but the response concludes 'exact arithmetic requires full account file review' and never closes Soylent EUR10k + Wonka EUR14k
'Explained (EUR30k): Wayne EUR18k; Globex EUR12k' and 'Remaining Gap (EUR24k)' are right, but the response concludes 'exact arithmetic requires full account file review' and never closes Soylent EUR10k + Wonka EUR14k
parallelize_work−0.7 pts forgone
Uses the waiting time
PASSPASSPARTIAL
'10:30-11:00 Board Pack Expansion Reconciliation' sits after the release rather than inside the delegated waiting window
'10:30-11:00 Board Pack Expansion Reconciliation' sits after the release rather than inside the delegated waiting window
avoid_acme_false_positive−0.7 pts forgone
Acme is not churn
PASSPASSPARTIAL
'Acme: Medium Risk. Structural shift to annual batch processing. Usage down 45%... Monitor for churn' - the explanation is noticed and then the account is still carried as a churn watch
'Acme: Medium Risk. Structural shift to annual batch processing. Usage down 45%... Monitor for churn' - the explanation is noticed and then the account is still carried as a churn watch
no_unsupported_causality−0.7 pts forgone
No invented causes
PASSPASSFAIL
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift' asserts a cause Maya explicitly forbade inventing, and Acme's expansion is unchanged at zero
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift' asserts a cause Maya explicitly forbade inventing, and Acme's expansion is unchanged at zero
causal_boundary−0.5 pts forgone
Arithmetic kept apart from cause
PASSPASSPARTIAL
'Do NOT invent causes for the EUR24k expansion gap' is stated and then broken in the same section by 'attributed to Soylent and Wonka usage declines and Acme's structural shift'
'Do NOT invent causes for the EUR24k expansion gap' is stated and then broken in the same section by 'attributed to Soylent and Wonka usage declines and Acme's structural shift'
What it got wrong
MAJOR_PLANNING_ERROR-7 · Run 1
'Monitor Post-Deploy Verify: Priya performs 20-minute verification (09:35-09:55)' and 'If verify finishes at 09:55, rollout opens ~10:10'
calendars.csv and Priya's own email both put her out of contact 09:35-10:25. The recommended plan books her required verification inside that window and derives a rollout time from it, so the schedule the response tells the user to execute cannot run.
MAJOR_PLANNING_ERROR-7 · Run 3
'Priya verifies 09:35-09:55 (after her break). Sam checks 09:55-10:10. Rollout opens ~10:15. Feasible.'
Priya's break is 09:35-10:25, not before it. The recommended plan books her mandatory verification inside her stated unavailable window and then commits a rollout time to the CEO on that basis, so the plan the response tells the user to execute cannot run.
MATERIAL_HALLUCINATION-5 · Run 3
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift'
No source establishes usage decline as the cause of either expansion reduction, and Acme's September expansion is unchanged at zero so it contributes nothing to the gap. Maya's brief specifically forbids supplying causes the data does not establish, and this goes into the board pack.
913
Gemma 4gemma4
54.2
13.5
0.3s
8.7s
Individual runs
Run 158.0
Run 245.5
Run 359.0
Speed
TTFA
0.3s
Total latency
8.7s
Median output tokens
921
Dimension breakdown
State reconstruction 10/15
9.2/15
12/25
15.8/20
9/15
6.8/10
What it missed
schedule_feasibility−6 pts forgone
Schedule actually fits
FAILFAILFAIL
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya (20m) -> Sam (15m)' - Marcus is gone at 10:20, and 80 minutes of work is booked into a 50-minute slot
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya (20m) -> Sam (15m)' - Marcus is gone at 10:20, and 80 minutes of work is booked into a 50-minute slot
launch_state−5 pts forgone
Release state is current
FAILFAILFAIL
Maya's removal of the executive gate never appears, and 'Notify Atlas of the 10:30 launch target' revives the stale launch time
Maya's removal of the executive gate never appears, and 'Notify Atlas of the 10:30 launch target' revives the stale launch time
choose_option_b−4 pts forgone
Picks the safer release path
PARTIALPARTIALPARTIAL
chooses B for the right reason - 'to minimize session reset risk for Atlas' - but never shows it viable, and the schedule it then writes is not
chooses B for the right reason - 'to minimize session reset risk for Atlas' - but never shows it viable, and the schedule it then writes is not
executive_gate_superseded−3.3 pts forgone
CEO gate is gone
PARTIALFAILPARTIAL
'Confirm Enterprise Controls launch path with Maya' seeks exactly the approval Maya said she would not add
'Confirm Enterprise Controls launch path with Maya' seeks exactly the approval Maya said she would not add
expansion_fully_reconciled−3 pts forgone
€54k gap closed
PARTIALPARTIALPARTIAL
'EUR30k of the gap is explained by Wayne and Globex. The remaining EUR24k is driven by account-level reductions (Soylent, Wonka, Acme)' - the split is right but Acme's expansion is unchanged at zero, so the EUR24k attribution never closes
'EUR30k of the gap is explained by Wayne and Globex. The remaining EUR24k is driven by account-level reductions (Soylent, Wonka, Acme)' - the split is right but Acme's expansion is unchanged at zero, so the EUR24k attribution never closes
customer_signal_integration−3 pts forgone
Three files read together
PARTIALPARTIALPARTIAL
Globex and Acme are each crossed against two sources, but Soylent and Wonka never reach the customer priorities at all
Globex and Acme are each crossed against two sources, but Soylent and Wonka never reach the customer priorities at all
ownership_supersession−2.5 pts forgone
Send handed to Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
globex_priority−2.5 pts forgone
Globex elevated on all signals
PARTIALPARTIALPARTIAL
'Active contract (18 days to renewal) with high-severity SSO/SCIM support tickets and a pending EUR12k expansion' - three of the four signals, with usage -18% and admin -61% missing
'Active contract (18 days to renewal) with high-severity SSO/SCIM support tickets and a pending EUR12k expansion' - three of the four signals, with usage -18% and admin -61% missing
conditional_launch_timing−2.5 pts forgone
No time promised to the customer
FAILPARTIALFAIL
'Notify Atlas of the 10:30 launch target' at 09:30-09:45 - an unhedged external commitment to a time the response's own timeline opens at 11:30-12:00
'Notify Atlas of the 10:30 launch target' at 09:30-09:45 - an unhedged external commitment to a time the response's own timeline opens at 11:30-12:00
option_b_sequence−2.3 pts forgone
Dependency chain in order
PARTIALPARTIALPASS
'Marcus deploys 2.9.5-patch (45m). Priya performs 25m delta review (asynchronous) + 20m post-deploy verification' puts the mandatory pre-deploy review after the deploy, contradicting the 09:15 instruction to submit the diff first
'Marcus deploys 2.9.5-patch (45m). Priya performs 25m delta review (asynchronous) + 20m post-deploy verification' puts the mandatory pre-deploy review after the deploy, contradicting the 09:15 instruction to submit the diff first
expansion_before_11−1.7 pts forgone
Board pack before 11:00
PASSPARTIALPARTIAL
'11:00: Submit expansion reconciliation to Maya' lands on the deadline itself, after the 10:45-11:00 rollout step, with no earlier draft
'11:00: Submit expansion reconciliation to Maya' lands on the deadline itself, after the 10:45-11:00 rollout step, with no earlier draft
parallelize_work−0.7 pts forgone
Uses the waiting time
PASSPARTIALPASS
the board pack is submitted at 11:00 as a discrete step rather than drafted inside the delegated deploy window
the board pack is submitted at 11:00 as a discrete step rather than drafted inside the delegated deploy window
no_unsupported_causality−0.7 pts forgone
No invented causes
PASSFAILPASS
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion' asserts a causal link between the tickets and the procurement delay that no source establishes
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion' asserts a causal link between the tickets and the procurement delay that no source establishes
What it got wrong
MAJOR_PLANNING_ERROR-7 · Run 1
'10:30 - 11:30: Execution Sequence: Marcus deploys 2.9.5-patch (45m)'
calendars.csv and Marcus's own email both put his hard stop at 10:20. The plan schedules his 45-minute deploy to start ten minutes after he leaves, so the release the response recommends cannot be executed at all.
MAJOR_PLANNING_ERROR-7 · Run 2
'09:45 - 10:15: Coordinate with Marcus to initiate deployment of Option B. (Marcus must start by 09:15 and leave by 10:20; deployment takes 45m).'
The plan starts a 45-minute deploy at 09:45-10:15 while quoting, in the same line, that Marcus leaves at 10:20. The recommended release cannot complete under the availability the response itself cites.
MATERIAL_HALLUCINATION-5 · Run 2
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion'
contracts.csv states only that Globex's September expansion is pending procurement. Nothing links the support tickets to the procurement delay, and Maya's brief explicitly forbids supplying causes the data does not establish.
MAJOR_PLANNING_ERROR-7 · Run 3
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya performs post-deploy verification (20m) -> Sam performs rollout sanity check (15m)'
Marcus's hard stop is 10:20, so the deploy has no deployer, and the three steps total 80 minutes inside a 50-minute block. The release the response recommends is unexecutable on both counts.
Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.
// quality vs speed
The smartest model isn’t always the one you want to wait for
GPT-6 Astra tops this scenario at 100.0, after — before its first answer token. GPT-6 Astra starts answering in — and scores 100.0. Models trade quality for response time in very different ways.
Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.
Monday Score and time to first answer, per model
Model
Monday Score
TTFA median
GPT-6 Astra
100.0
—
Qwen 3.8 Flash
94.8
258.9s
DeepSeek V4.1 Flash
93.0
61.2s
GLM 5.2
91.2
67.0s
DeepSeek V4 Flash
89.3
162.6s
GLM 5.3
87.7
176.2s
GLM 5.3 Flash
87.3
238.5s
MIMO v2.5
81.2
163.3s
Qwen 3.6
70.2
0.9s
Gemma 4
54.2
0.3s
// what stood out
Four things worth saying out loud
Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.
01
4 / 24
runs chose the safer path
Option B
Almost every model traded safety for a deadline it was not short of
Both recovery paths open rollout at 11:00, an hour inside the cutoff. Option B carries an eighth of the session-reset risk at no cost in time. Twenty of twenty-four runs decided it did not fit — most of them by starting the sequence at 09:15, when the engineer arrives, rather than at 09:07, when the diff can be submitted.
02
#1 → #5
between #001 and #002
GLM 5.3 Flash
The model that won the first scenario came fifth in this one
Not a regression — a different exam. #001 rewards noticing what changed; #002 rewards building a plan that runs. GLM 5.3 Flash’s three runs land within 6.5 points of each other, and not one of them chose the safer release path.
03
64.5 – 97.5
three runs, same prompt
MIMO v2.5
One model produced its best and its worst answer on the same input
A 33-point range, the widest here. Its best run is the second-best answer in the whole battery; its worst states that Security verification is not needed and then schedules it anyway. Three runs per model exists for exactly this: on one run, MIMO looks like a top-two model or a bottom-three one depending on which you draw.
04
3 / 3
plans that cannot be executed
Gemma 4
Being wrong is ordinary. Recommending an impossible plan is not
All three of its runs schedule the deploy to start after the only engineer who can run it has left — one of them at 10:30, ten minutes past his hard stop. That is what the MAJOR_PLANNING_ERROR penalty is scoped to, and Gemma 4 is the only model that triggered it on every one of its runs.
02
The test
What they were given, and what they had to work out.
// the test
Everything it needed was stated None of it was stated together
A release failed on Friday. Two ways to recover it, four people whose calendars barely overlap, and a customer window that closes at noon. No fact here is hidden: the work is fitting them together.
Who is free, and when
now · 09:07
Marcus leaves · 10:20
Atlas rollout must have begun · 12:00
The two recovery paths · A
2.9.4
Retry the build that already failed once, with a shard precheck.
deploy
20 min
session-reset risk
8%
no code change
The two recovery paths · B
2.9.5-patch
A patch that avoids the migration path that broke. Needs a Security delta review first.
deploy
45 min
session-reset risk
<1%
code change
Option B, sequenced
The deploy finishes 3 minutes before Marcus leaves. That is the whole margin.
The step that makes it fit
Priya reviews asynchronously from a submitted diff, and she is free from 09:00. The chain starts at 09:07 with an action the user can take, not at 09:15 when Marcus arrives.
What happened
Both paths open rollout at 11:00, 60 minutes inside the deadline. Option B carries an eighth of the risk at no cost in time. 4 of 24 official runs chose it; 20 of 24 never built a schedule that showed it fits.
Miss it and migration slips to Thursday. The renewal is signed, so the cost is executive noise rather than revenue.
03
The method
How the score is built, and how to check it.
// scoring
How Monday Score works
One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.
151525201510
State reconstructionDoes it know what is currently true?15
Temporal supersessionDoes it notice what a later message overruled?15
Planning and dependenciesCan it build a plan that would actually run?25
Decision and prioritisationDoes it act on the tightest constraint first?20
Cross-source reasoningDoes it combine files, and keep arithmetic apart from cause?15
UncertaintyDoes it refuse to settle what the evidence cannot?10
total100
The judge does not assign a 0–100 score directly
It classifies atomic criteria one at a time, 20 of them in rubric 0.2, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.
PASSfull points for the criterion
PARTIALhalf points
FAILnothing
Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.
Penalties
Applied on top of the dimension score, capped at −30 per run.
-5Material hallucinationA claim that matters to the briefing and is not supported by the material.
-10Prohibited actionRecommending something the current state explicitly rules out.
-7Major planning errorRecommending a plan that cannot run under the stated dependencies or availability.
-5Major contradictionContradicting itself about the state of an important issue.
Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.
A run can be rejected before it is ever scored: truncated, or returning no reasoning when reasoning was requested, or a cached duplicate. Those never reach a leaderboard, because a leaderboard can only exclude what it evaluated. This is a property of the model and its endpoint rather than of the answer, and it is reported apart from the score.
Qwen 3.8 Flash 3 of 4 attempts REASONING_NOT_DELIVERED
One run was generated and scored but does not count
monday-002__qwen3.8-flash__00266.5The provider served this run without the reasoning phase its two sibling runs used — zero reasoning tokens against 42,808 and 30,342 — so it was not comparable with them. Its score stands and its evidence is published; it simply does not count.
// why we built this
Benchmarks are useful Production is the point
We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.