48 Slack messages · 7 emails · 2 meetingsthree of them change what is already true
metrics.csv Umbrella’s anomaly — in the metrics, mentioned by nobody
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models onlyAll models
monday-001 · 13 models, 3 runs each
#
Model
Monday Score
Run range
TTFA median
Total latency
Show detail for
1
GPT-6 Astragpt-6-astraclosed
100.0
0.0
—
—
Individual runs
Run 1100.0
Run 2100.0
Run 3100.0
Speed
TTFA
—
Total latency
—
Median output tokens
—
Dimension breakdown
Action detection 25/25
State reconstruction 25/25
Prioritization 15/15
Data reasoning 15/15
Uncertainty 10/10
Noise resistance 10/10
What it missed
Nothing. Full marks on every criterion.
What it got wrong
No penalties and no false actions in any run.
2
Claude Opus 5closed
98.5
0.0
–
–
13
GLM 5.3 Flashglm5.3-flash
98.0
1.5
118.8s
132.4s
Individual runs
Run 198.5
Run 297.0
Run 398.5
Speed
TTFA
118.8s
Total latency
132.4s
Median output tokens
5,603
Dimension breakdown
Action detection 25/25
State reconstruction 25/25
Prioritization 15/15
Data reasoning 13.5/15
Uncertainty 9.5/10
Noise resistance 10/10
What it missed
umbrella_error_rate−1.5 pts forgone
Errors read against traffic
PARTIALPARTIALPARTIAL
'1821ms (vs ~300 for everyone else)' quotes the peers' latency and 'Stark/Contoso error rates negligible' reaches for the rate idea, but no request volume and no rate figure - the 163-against-94,217 reading is never made
'1821ms (vs ~300 for everyone else)' quotes the peers' latency and 'Stark/Contoso error rates negligible' reaches for the rate idea, but no request volume and no rate figure - the 163-against-94,217 reading is never made
friday_conversion_unknown−0.5 pts forgone
Friday conversion unverifiable
PASSPARTIALPASS
'Julia saw weak conversion Friday but suspects noisy traffic; no owner assigned' - flagged, but no statement that the provided data cannot settle it
'Julia saw weak conversion Friday but suspects noisy traffic; no owner assigned' - flagged, but no statement that the provided data cannot settle it
What it got wrong
No penalties and no false actions in any run.
4
Fable 5.1closed
97.5
1.5
–
–
25
GLM 5.3glm5.3
97.5
6.0
181.6s
202.8s
Individual runs
Run 1100.0
Run 294.0
Run 398.5
Speed
TTFA
181.6s
Total latency
202.8s
Median output tokens
6,861
Dimension breakdown
Action detection 23.8/25
State reconstruction 25/25
Prioritization 14.2/15
Data reasoning 14.5/15
Uncertainty 10/10
Noise resistance 10/10
What it missed
umbrella_action−1.2 pts forgone
Umbrella action
PASSPARTIALPASS
detected and actioned for today ('Raise with engineering/CS today'), no root cause claimed - but ranked #4, below a holding reply to Initech and below the churn analysis, so it is under-prioritised; between PASS and PARTIAL the rubric's tie-break takes the lower
detected and actioned for today ('Raise with engineering/CS today'), no root cause claimed - but ranked #4, below a holding reply to Initech and below the churn analysis, so it is under-prioritised; between PASS and PARTIAL the rubric's tie-break takes the lower
umbrella_and_churn_high−0.8 pts forgone
Umbrella and churn prioritised
PASSPARTIALPASS
churn is appropriately placed at #3; Umbrella is at #4 behind a routine email reply, so only one of the two is properly prioritised
churn is appropriately placed at #3; Umbrella is at #4 behind a routine email reply, so only one of the two is properly prioritised
umbrella_error_rate−0.5 pts forgone
Errors read against traffic
PASSPASSPARTIAL
compares Umbrella with the other accounts on latency ('every other account: 267-341 ms') and notes they look healthy, but never reads the 163 errors against Umbrella's 94,217 requests, gives no error rate, and does not note that Stark carries more raw errors on far higher traffic
compares Umbrella with the other accounts on latency ('every other account: 267-341 ms') and notes they look healthy, but never reads the 163 errors against Umbrella's 94,217 requests, gives no error rate, and does not note that Stark carries more raw errors on far higher traffic
What it got wrong
No penalties and no false actions in any run.
6
Gemini 3.8 Flash (High)closed
97.0
0.0
–
13.2s
37
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
95.0
6.0
39.3s
44.1s
Individual runs
Run 191.0
Run 297.0
Run 397.0
Speed
TTFA
39.3s
Total latency
44.1s
Median output tokens
8,154
Dimension breakdown
Action detection 23.8/25
State reconstruction 25/25
Prioritization 14.2/15
Data reasoning 12/15
Uncertainty 10/10
Noise resistance 10/10
What it missed
umbrella_error_rate−3 pts forgone
Errors read against traffic
FAILFAILFAIL
The response reports Umbrella's anomaly but does not contextualize the 163 errors against request volume or another account.
No usable error-rate comparison is made.
umbrella_action−1.2 pts forgone
Umbrella action
PARTIALPASSPASS
Umbrella is noticed, but the investigation is not promoted strongly enough into today's action list.
The response recognizes the Umbrella anomaly but gives it weaker operational treatment than the PASS band requires.
umbrella_and_churn_high−0.8 pts forgone
Umbrella and churn prioritised
PARTIALPASSPASS
Churn is actionable, while Umbrella is not kept at the same high-priority level.
Only one of the two required high-priority investigations is consistently elevated.
What it got wrong
No penalties and no false actions in any run.
48
Qwen 3.8 Flashqwen3.8-flash
92.5
6.0
72.2s
92.0s
Individual runs
Run 193.5
Run 295.0
Run 389.0
Speed
TTFA
72.2s
Total latency
92.0s
Median output tokens
9,253
Dimension breakdown
Action detection 25/25
State reconstruction 25/25
Prioritization 11.5/15
Data reasoning 12.5/15
Uncertainty 8.5/10
Noise resistance 10/10
What it missed
no_false_priorities−2.7 pts forgone
No false priorities
PARTIALPARTIALFAIL
two cleanup items outrank the churn analysis: the resolved Acme incident follow-up at #3 and the Analytics v2 task cleanup at #4, which closes the launch-email task Julia already owns
two cleanup items outrank the churn analysis: the resolved Acme incident follow-up at #3 and the Analytics v2 task cleanup at #4, which closes the launch-email task Julia already owns
umbrella_error_rate−2.5 pts forgone
Errors read against traffic
FAILPARTIALFAIL
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
'The team suspected noisy traffic, but it has not been investigated' - flagged, but no statement that the provided data cannot settle it
'The team suspected noisy traffic, but it has not been investigated' - flagged, but no statement that the provided data cannot settle it
umbrella_and_churn_high−0.8 pts forgone
Umbrella and churn prioritised
PASSPASSPARTIAL
Umbrella is well placed at #2, but the churn analysis is #6 of eight, below two task-cleanup items and the Initech reply
Umbrella is well placed at #2, but the churn analysis is #6 of eight, below two task-cleanup items and the Initech reply
What it got wrong
No penalties and no false actions in any run.
59
DeepSeek V4 Flashdeepseek-v4-flash
90.8
8.0
55.4s
63.3s
Individual runs
Run 195.5
Run 287.5
Run 389.5
Speed
TTFA
55.4s
Total latency
63.3s
Median output tokens
4,871
Dimension breakdown
Action detection 22.7/25
State reconstruction 25/25
Prioritization 12.7/15
Data reasoning 12/15
Uncertainty 8.5/10
Noise resistance 10/10
What it missed
umbrella_error_rate−3 pts forgone
Errors read against traffic
FAILFAILFAIL
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
umbrella_action−2.3 pts forgone
Umbrella action
PASSPARTIALPARTIAL
only under 'unclear': 'Umbrella metrics look anomalous... Needs an explanation' - never enters today's list
only under 'unclear': 'Umbrella metrics look anomalous... Needs an explanation' - never enters today's list
umbrella_and_churn_high−1.7 pts forgone
Umbrella and churn prioritised
PASSPARTIALPARTIAL
churn actionable at #4; Umbrella appears only under 'unclear'
churn actionable at #4; Umbrella appears only under 'unclear'
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
'possibly noise, not yet investigated'; 'real or noise? Needs a quick check' - flagged, but no statement that the provided data cannot settle it
'possibly noise, not yet investigated'; 'real or noise? Needs a quick check' - flagged, but no statement that the provided data cannot settle it
no_false_priorities−0.7 pts forgone
No false priorities
PASSPARTIALPASS
the #3 task-cleanup bundle, which closes the resolved Acme task, is ranked above the churn analysis
the #3 task-cleanup bundle, which closes the resolved Acme task, is ranked above the churn analysis
What it got wrong
No penalties and no false actions in any run.
610
MIMO v2.5mimo-v2.5
88.2
6.5
42.4s
64.1s
Individual runs
Run 191.0
Run 289.0
Run 384.5
Speed
TTFA
42.4s
Total latency
64.1s
Median output tokens
2,150
Dimension breakdown
Action detection 20.5/25
State reconstruction 25/25
Prioritization 11.2/15
Data reasoning 13/15
Uncertainty 8.5/10
Noise resistance 10/10
What it missed
umbrella_action−3.5 pts forgone
Umbrella action
PARTIALPARTIALPARTIAL
only under 'Unclear': 'Worth investigating - is there an active issue, or was this just a quiet week?' - never enters today's list
only under 'Unclear': 'Worth investigating - is there an active issue, or was this just a quiet week?' - never enters today's list
umbrella_and_churn_high−2.5 pts forgone
Umbrella and churn prioritised
PARTIALPARTIALPARTIAL
churn actionable in today's list; Umbrella appears only under 'Unclear'
churn actionable in today's list; Umbrella appears only under 'Unclear'
umbrella_error_rate−2 pts forgone
Errors read against traffic
PARTIALPARTIALFAIL
the three figures are given in isolation: no request volume, no rate, and no reference to any other customer
the three figures are given in isolation: no request volume, no rate, and no reference to any other customer
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
'flagged as potentially noisy Friday traffic. Worth revisiting next week' - not repeated as fact, but no data-gap statement
'flagged as potentially noisy Friday traffic. Worth revisiting next week' - not repeated as fact, but no data-gap statement
no_false_priorities−1.3 pts forgone
No false priorities
PASSPARTIALPARTIAL
Globex, which the source says has no urgency, is #4, above the Umbrella investigation
Globex, which the source says has no urgency, is #4, above the Umbrella investigation
churn_action−1 pts forgone
Churn action
PASSPASSPARTIAL
the board deadline is stated ('Need a fuller breakdown for Tuesday's board meeting'), but the analysis never enters the action table - it is filed under 'no action required from you' as 'Analysis task is yours for Tuesday'
the board deadline is stated ('Need a fuller breakdown for Tuesday's board meeting'), but the analysis never enters the action table - it is filed under 'no action required from you' as 'Analysis task is yours for Tuesday'
What it got wrong
No penalties and no false actions in any run.
711
GLM 5.2glm5.2
84.3
16.0
2.7s
6.9s
Individual runs
Run 187.0
Run 275.0
Run 391.0
Speed
TTFA
2.7s
Total latency
6.9s
Median output tokens
1,241
Dimension breakdown
Action detection 20.3/25
State reconstruction 25/25
Prioritization 11.2/15
Data reasoning 10/15
Uncertainty 7.8/10
Noise resistance 10/10
What it missed
umbrella_action−4.7 pts forgone
Umbrella action
PARTIALFAILPARTIAL
Umbrella is never mentioned; the CSV is opened only to check that Initech is healthy
Umbrella is never mentioned; the CSV is opened only to check that Initech is healthy
umbrella_and_churn_high−2.5 pts forgone
Umbrella and churn prioritised
PARTIALPARTIALPARTIAL
churn well placed at #3, but the Umbrella investigation is buried at #6
churn well placed at #3, but the Umbrella investigation is buried at #6
umbrella_error_rate−2 pts forgone
Errors read against traffic
PARTIALFAILPARTIAL
no Umbrella reasoning at all
no Umbrella reasoning at all
umbrella_detected−1.7 pts forgone
Umbrella anomaly detected
PASSFAILPASS
the anomaly is missed entirely - the only customer row it reads is Initech's, and it reads it correctly as healthy
the anomaly is missed entirely - the only customer row it reads is Initech's, and it reads it correctly as healthy
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
'Is it noise or a real signal? Needs a look at the dashboard' - uncertainty flagged, but no statement that the provided data cannot settle it
'Is it noise or a real signal? Needs a look at the dashboard' - uncertainty flagged, but no statement that the provided data cannot settle it
no_false_priorities−1.3 pts forgone
No false priorities
FAILPASSPASS
two false priorities outrank the Umbrella investigation: the resolved Acme task close-out at #4 and Globex at #5, which the source says has no urgency
two false priorities outrank the Umbrella investigation: the resolved Acme task close-out at #4 and Globex at #5, which the source says has no urgency
umbrella_metrics−1.3 pts forgone
Umbrella metrics cited
PASSFAILPASS
no Umbrella figures
no Umbrella figures
umbrella_unknown_cause−0.7 pts forgone
Umbrella cause left open
PASSPARTIALPASS
invents no cause, but expresses no uncertainty either - Umbrella is absent
invents no cause, but expresses no uncertainty either - Umbrella is absent
What it got wrong
No penalties and no false actions in any run.
812
Qwen 3.6qwen3.6
82.8
5.0
0.8s
6.2s
Individual runs
Run 182.5
Run 280.5
Run 385.5
Speed
TTFA
0.8s
Total latency
6.2s
Median output tokens
1,036
Dimension breakdown
Action detection 21.5/25
State reconstruction 25/25
Prioritization 11.2/15
Data reasoning 12/15
Uncertainty 6.5/10
Noise resistance 10/10
What it missed
umbrella_action−3.5 pts forgone
Umbrella action
PARTIALPARTIALPARTIAL
under 'does not require action': 'a potential churn risk, though not explicitly flagged as urgent... Worth keeping in mind' - noticed, but not work for today
under 'does not require action': 'a potential churn risk, though not explicitly flagged as urgent... Worth keeping in mind' - noticed, but not work for today
umbrella_error_rate−3 pts forgone
Errors read against traffic
FAILFAILFAIL
the 163 errors are never mentioned; no volume, no rate, no other customer's figures
the 163 errors are never mentioned; no volume, no rate, no other customer's figures
umbrella_and_churn_high−2.5 pts forgone
Umbrella and churn prioritised
PARTIALPARTIALPARTIAL
churn actionable at #2 HIGH; Umbrella filed under 'does not require action'
churn actionable at #2 HIGH; Umbrella filed under 'does not require action'
umbrella_unknown_cause−2 pts forgone
Umbrella cause left open
PARTIALPARTIALPARTIAL
no cause asserted as established, but no uncertainty expressed either, and Umbrella is absent from section 4
no cause asserted as established, but no uncertainty expressed either, and Umbrella is absent from section 4
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
the Friday conversion question is never raised: the claim is not repeated, but the gap is not flagged
the Friday conversion question is never raised: the claim is not repeated, but the gap is not flagged
no_false_priorities−1.3 pts forgone
No false priorities
PASSPARTIALPARTIAL
Globex, which the source says has no urgency, is #2 High, above the churn analysis
Globex, which the source says has no urgency, is #2 High, above the churn analysis
What it got wrong
MATERIAL_HALLUCINATION-5 · Run 1
Strictly confidential until Wednesday's review / state that new tiers are under review and will be announced publicly soon
The source says pricing stays confidential until approved AND publicly announced; Wednesday is only the next review, and no announcement is scheduled at all. The error is written into the reply the user is told to send Initech, so it leaves the company.
MATERIAL_HALLUCINATION-5 · Run 2
Strictly confidential until Wednesday's review / state that pricing is under review and will be communicated officially soon
The source says pricing stays confidential until approved AND publicly announced; Wednesday is only the next review. A reader could conclude the tiers may be shared after Wednesday.
913
Gemma 4gemma4
70.7
6.5
0.6s
5.5s
Individual runs
Run 168.5
Run 268.5
Run 375.0
Speed
TTFA
0.6s
Total latency
5.5s
Median output tokens
605
Dimension breakdown
Action detection 15.3/25
State reconstruction 23.3/25
Prioritization 12.5/15
Data reasoning 3/15
Uncertainty 6.5/10
Noise resistance 10/10
What it missed
umbrella_action−7 pts forgone
Umbrella action
FAILFAILFAIL
Umbrella never mentioned; the metrics file is not used at all
Umbrella never mentioned; the metrics file is not used at all
umbrella_detected−5 pts forgone
Umbrella anomaly detected
FAILFAILFAIL
anomaly missed
anomaly missed
umbrella_metrics−4 pts forgone
Umbrella metrics cited
FAILFAILFAIL
no supporting figures at all
no supporting figures at all
umbrella_error_rate−3 pts forgone
Errors read against traffic
FAILFAILFAIL
no metrics reasoning of any kind
no metrics reasoning of any kind
contoso_action−2.7 pts forgone
Contoso action
PARTIALPARTIALPASS
first item under High with the 11:00 deadline, but 'ready to renew pending a single legal clarification' states sequence, not jeopardy: nothing says what happens if 11:00 is missed
first item under High with the 11:00 deadline, but 'ready to renew pending a single legal clarification' states sequence, not jeopardy: nothing says what happens if 11:00 is missed
umbrella_and_churn_high−2.5 pts forgone
Umbrella and churn prioritised
PARTIALPARTIALPARTIAL
churn is a High item; Umbrella is entirely absent
churn is a High item; Umbrella is entirely absent
umbrella_unknown_cause−2 pts forgone
Umbrella cause left open
PARTIALPARTIALPARTIAL
invents no cause, but expresses no uncertainty either - Umbrella is absent
invents no cause, but expresses no uncertainty either - Umbrella is absent
contoso_blocked−1.7 pts forgone
Renewal blocked
PARTIALPARTIALPASS
'pending a single legal clarification' - identifies the question, misses the signature or approval consequence
'pending a single legal clarification' - identifies the question, misses the signature or approval consequence
friday_conversion_unknown−1.5 pts forgone
Friday conversion unverifiable
PARTIALPARTIALPARTIAL
'currently unconfirmed if this was a data anomaly or a trend' - uncertainty noted, data gap not flagged
'currently unconfirmed if this was a data anomaly or a trend' - uncertainty noted, data gap not flagged
What it got wrong
No penalties and no false actions in any run.
Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.
// quality vs speed
The smartest model isn’t always the one you want to wait for
GPT-6 Astra tops this scenario at 100.0, after — before its first answer token. GPT-6 Astra starts answering in — and scores 100.0. Models trade quality for response time in very different ways.
Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.
Monday Score and time to first answer, per model
Model
Monday Score
TTFA median
GPT-6 Astra
100.0
—
GLM 5.3 Flash
98.0
118.8s
GLM 5.3
97.5
181.6s
DeepSeek V4.1 Flash
95.0
39.3s
Qwen 3.8 Flash
92.5
72.2s
DeepSeek V4 Flash
90.8
55.4s
MIMO v2.5
88.2
42.4s
GLM 5.2
84.3
2.7s
Qwen 3.6
82.8
0.8s
Gemma 4
70.7
0.6s
// what actually happened
Four things the results showed
Eight models, twenty-four runs, one judge. These are the patterns worth reporting, not the ones that make the best headline.
01
GLM 5.3 Flash almost solved Monday
It was not only the highest-scoring model. It was also the most consistent: three runs inside a point and a half of each other, with full marks on action detection, state reconstruction and prioritization.
Where it lost points It spotted the Umbrella anomaly but never fully reasoned about errors relative to request volume.
98.0/100
GLM 5.3 Flash
run 198.5
run 297.0
run 398.5
range1.5
Perfect on three of six dimensions
02
One run isn’t enough
A single run could make GLM 5.2 look like a 75-point model or a 91-point model. That is why every official MondayBench result uses three runs, and why the leaderboard reports the spread next to the score instead of hiding it inside a mean.
DeepSeek V4 Flash
run 1: 95.5run 2: 87.5run 3: 89.5
95.5 / 87.5 / 89.5 range 8.0
GLM 5.2
run 1: 87.0run 2: 75.0run 3: 91.0
87.0 / 75.0 / 91.0 range 16.0
65100
03
Some models read the conversation but miss the business
Gemma understood much of the conversation, but missed the most important signal hidden in the customer metrics. Umbrella was never mentioned in any of its three runs: the metrics file was effectively ignored.
Data reasoning State reconstruction stayed high: it followed the thread, it just never opened the spreadsheet.
3/15
Data reasoning
Gemma 470.7
runs naming Umbrella0/3
04
Fast doesn’t mean safe
Qwen 3.6 was one of the fastest models in the benchmark, but it was also the only model to receive material hallucination penalties.
The scenario states that proposed pricing stays confidential until it is approved and publicly announced. In two of three runs Qwen 3.6 rewrote that as confidentiality ending after Wednesday’s review, or as an announcement being imminent, and put it in the reply it told the user to send the customer. Its third run states the rule correctly.
2
penalties
ttfa0.8s
total6.2s
score82.8
02
The test
What they were given, and what they had to work out.
// the input
This isn’t a trivia test It’s a state reconstruction problem
Real work is messy. Information arrives through different channels, tasks become stale, decisions change and the most important signal may never be explicitly flagged.
slack
#incidents Thu 09:48 · Marcus Webb
Rollback complete. Error rates are returning to normal.
#sales Fri 08:51 · Noah Williams
Contoso update: call went really well. Commercial terms are basically agreed.
#product Fri 09:36 · Sam Rivera
The export regression is more annoying than we thought. I’m proposing Wednesday instead of Monday.
#leadership Fri 11:42 · Daniel Foster
The dashboard has it moving from 3.1% to 4.4%. Would be good to understand what’s driving that before the board meeting.
tasks
Priority
Task
Owner
Due
Urgent
Investigate Acme API failures stale
You
Thursday
High
Confirm Contoso renewal next steps
You
Monday
High
Prepare Analytics v2 launch email stale
You
Monday
Medium
Analyse enterprise churn
You
Tuesday
inbox
Fri 16:18 · sarah@contoso.example
One remaining point before renewal
…our legal team flagged one remaining question regarding data retention. Could you
confirm this for us by 11:00 Monday?
If we can’t get legal comfortable with this point before then, we’ll need to push the
renewal signature into our next approval window.
metrics · last 24h
Customer
Requests
Errors
p95
7d
Acme
482,140
8
312ms
+3%
Contoso
821,334
22
284ms
+7%
Umbrella
94,217
163
1,821ms
−41%
Stark Industries
1,482,012
74
298ms
+9%
Everything the models saw, in the order they saw it. The full scenario is frozen and published with the results.
// the answer key
What the model needed to figure out
Five judgements decide most of the score. They are published with the scenario, so you can decide for yourself whether the benchmark is asking the right things.
threadwhat the sources saycorrect conclusion
01
Contoso
€72k ARR renewal.
Legal blocker.
Written technical confirmation required before 11:00.
Verify the backup-deletion behavior with the technical owner and respond before 11:00.
02
Acme
Task list still says urgent.
The incident is already resolved.
Do not treat it as an active incident.
03
Analytics v2
Task list says Monday launch.
Latest information says Wednesday.
The launch email is owned by Julia.
Don’t execute the stale launch tasks.
04
Umbrella
Nobody mentions it in Slack or email.
The anomaly only exists in the metrics.
Investigate it today, but don’t invent a cause.
05
Churn
Enterprise churn rose from 3.1% to 4.4%.
There are known cancellations but insufficient evidence to explain the full increase.
Investigate the breakdown before Tuesday’s board meeting without claiming causality.
03
The method
How the score is built, and how to check it.
// scoring
How Monday Score works
One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.
252515151010
Action detectionDid it find the work that had to happen today?25
State reconstructionDoes it know what is currently true?25
PrioritizationIs the order right, and is nothing stale promoted?15
Data reasoningDid it read the numbers, and read them correctly?15
UncertaintyDoes it separate what it knows from what it assumes?10
Noise resistanceDoes it stay away from work that is already done?10
total100
The judge does not assign a 0–100 score directly
It classifies atomic criteria one at a time, 19 of them in rubric 0.2, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.
PASSfull points for the criterion
PARTIALhalf points
FAILnothing
Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.
Penalties
Applied on top of the dimension score, capped at −20 per run.
-5Material hallucinationA claim that matters to the briefing and is not supported by the material.
-10Prohibited actionRecommending something the current state explicitly rules out.
-5Major contradictionContradicting itself about the state of an important issue.
Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.