New: DeepSeek V4.1 Flash.The architecture built for agents, and the 890 bytes that change an agent bill.Explore the piece →
// mondaybench #004beta
We gave 10 open models the same messy Monday morning
Eight files, seven of them tables. Two months of a SaaS company’s revenue, usage, churn, acquisition and support — and three decisions waiting on numbers that none of the tables contains.
customers.csv 36 rows customer_idsegment, plan, ARR, renewal, status
revenue.csv 71 rows customer_idinvoices, two months, two sources
usage.csv 67 rows usage_account_idrequests, tokens, active users
plans.csv 6 rows planrevenue, infra and support cost, included volume
acquisition.csv 15 rows customer_idchannel, new logo or expansion, attributed ARR
funnel.csv 8 rows channelleads and paid trials, by channel and month
support.csv 133 rows customer_idtickets, severity, status
usage_account_idjoins on a different key — two accounts spell it another way
01
The result
Ten models, 30 runs, one judge.
// results
The leaderboard
Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.
Open models onlyAll models
monday-001 · 13 models, 3 runs each
#
Model
Monday Score
Run range
TTFA median
Total latency
Show detail for
11
GLM 5.3 Flashglm5.3-flash
98.0
1.5
766.2s
827.8s
Individual runs
Run 198.5
Run 298.5
Run 397.0
Speed
TTFA
766.2s
Total latency
827.8s
Median output tokens
36,248
Dimension breakdown
20/20
15/15
19/19
19/19
16.5/18
8.5/9
What it missed
figures_carried_correctly−1.5 pts forgone
Carries its own figures through correctly
PARTIALPARTIALPARTIAL
'wonka ... 7 tickets in October - 5 of them rate_limits' when 4 of the 7 are rate_limits; and 'wayne renews in ~5 weeks' when 2026-11-30 is four weeks after the review date
Two small slips in supporting detail. Everything else checkable is exact, including the EUR 262.9k -> EUR 246.5k profit walk, the +5.8% cost growth, the MRR bridge, the EUR 216k expansion bookings itemised, the EUR 3.23M of ARR with open high-severity tickets, all eight funnel cells, the full new-logo table, all twelve plan figures, contribution per customer for all three plans, and the mid-market -EUR 42k walk.
segment_recommendation−0.5 pts forgone
Does not recommend cutting the segment
PASSPASSPARTIAL
'Yes - reduce acquisition investment in SMB, redirect rather than cut ... pause incremental paid SMB acquisition, redirect budget to SMB retention/onboarding ... Revisit once CAC-by-segment exists.'
A reduction hedged with conditions, which is the PARTIAL band precisely. It names the missing input the PASS bar asks for and the budget stays inside SMB rather than leaving it, but leadership is still told to pause acquisition spend.
What it got wrong
No penalties and no false actions in any run.
2
GPT-6 Astragpt-6-astraclosed
98.0
0.0
—
—
Individual runs
Run 198.0
Run 298.0
Run 398.0
Speed
TTFA
—
Total latency
—
Median output tokens
—
Dimension breakdown
20/20
15/15
19/19
19/19
16/18
9/9
What it missed
support_normalized−2 pts forgone
Compares support load on a rate, not a raw count
PARTIALPARTIALPARTIAL
PARTIAL: Pro's ticket movement is quantified (19 -> 28, rate limits 5 -> 13) and support cost sits per plan in the contribution table, but no rate is computed on any denominator and Enterprise's raw share of tickets is never addressed
PARTIAL: Pro's ticket movement is quantified (19 -> 28, rate limits 5 -> 13) and support cost sits per plan in the contribution table, but no rate is computed on any denominator and Enterprise's raw share of tickets is never addressed
What it got wrong
No penalties and no false actions in any run.
23
Qwen 3.8 Flashqwen3.8-flash
96.5
3.0
394.1s
468.0s
Individual runs
Run 196.5
Run 298.0
Run 395.0
Speed
TTFA
394.1s
Total latency
468.0s
Median output tokens
52,295
Dimension breakdown
20/20
15/15
19/19
19/19
15/18
8.5/9
What it missed
support_normalized−2 pts forgone
Compares support load on a rate, not a raw count
PARTIALPARTIALPARTIAL
'October support tickets by plan: Enterprise 34 -> 35, Pro 19 -> 28, Starter 7 -> 11' and 'September Enterprise-segment tickets: 26, October: 25'
It avoids the trap - Enterprise is called stable and Pro gets the focus - and it reads the segment view as well as the plan view, which is more care than most. But no rate against any denominator is computed, so the normalisation the criterion asks for is not performed.
figures_carried_correctly−1 pts forgone
Carries its own figures through correctly
PARTIALPASSPARTIAL
'Enterprise 34 September tickets' where the enterprise plan has 33, which also makes its own column sum to 60 against a September total of 59
One miscount. Everything else checkable is exact, and the response carries more arithmetic than any other in the battery: the ARR bridge reconciling to EUR 8,788,800 to the euro, four revenue bases, the segment ARR walk including the reconstructed EUR 1,704,000, both channel tables with totals, the full two-month plan economics with cost subtotals, the utilisation table computed from customer counts, all five over-allowance accounts with overage volumes, the Soylent four-metric table, and the 13 October rate-limit tickets itemised per account - the only response anywhere to get that count right.
segment_recommendation−0.5 pts forgone
Does not recommend cutting the segment
PASSPASSPARTIAL
'the data do not support a full cut without CAC and segment cost data' and 'Reallocate SMB investment toward onboarding, usage activation, and retention measurement rather than cutting the segment entirely' - but also 'For SMB, reduce or pause low-quality acquisition spend, especially starter acquisition, until activation and retention improve.'
A reduction hedged with conditions, which is the PARTIAL band. It names the missing inputs the PASS bar asks for and keeps the budget inside SMB, but leadership is still told to pause acquisition spend rather than to investigate first.
What it got wrong
No penalties and no false actions in any run.
34
GLM 5.3glm5.3
91.7
17.0
634.2s
673.2s
Individual runs
Run 195.0
Run 281.5
Run 398.5
Speed
TTFA
634.2s
Total latency
673.2s
Median output tokens
46,592
Dimension breakdown
19/20
14.2/15
19/19
17.3/19
13.7/18
8.5/9
What it missed
figures_carried_correctly−2 pts forgone
Carries its own figures through correctly
PARTIALFAILPARTIAL
five incorrect figures across different parts of the analysis, one of them a headline that contradicts its own components. The ARR bridge reads 'EUR8,953k to EUR8,789k = -EUR164k (-1.8%)' while the components listed beneath it (+246, +133, +36, -30, -304) sum to +EUR81k; the closing figure is right, the opening figure is unreachable, and the stated direction is wrong - ARR on 1 September was about EUR8,713k, so it rose. Also: Acme is put at '84%' of its 3M included requests when 2.31M is 77%; '4 of 12 September SMB logos' where there are 13; rate-limit tickets 'tripled (4 to 13)' where September is 5; and 'infra per token ~EUR11.7/M' which divides the enterprise PLAN's infra cost by the enterprise SEGMENT's tokens, the plan-consistent figure being about EUR9.7/M - a wrong secondary rate. Repeated incorrect figures across different parts, plus one materially wrong: the FAIL band
five incorrect figures across different parts of the analysis, one of them a headline that contradicts its own components. The ARR bridge reads 'EUR8,953k to EUR8,789k = -EUR164k (-1.8%)' while the components listed beneath it (+246, +133, +36, -30, -304) sum to +EUR81k; the closing figure is right, the opening figure is unreachable, and the stated direction is wrong - ARR on 1 September was about EUR8,713k, so it rose. Also: Acme is put at '84%' of its 3M included requests when 2.31M is 77%; '4 of 12 September SMB logos' where there are 13; rate-limit tickets 'tripled (4 to 13)' where September is 5; and 'infra per token ~EUR11.7/M' which divides the enterprise PLAN's infra cost by the enterprise SEGMENT's tokens, the plan-consistent figure being about EUR9.7/M - a wrong secondary rate. Repeated incorrect figures across different parts, plus one materially wrong: the FAIL band
support_normalized−1.3 pts forgone
Compares support load on a rate, not a raw count
PARTIALPARTIALPASS
support cost as a share of plan revenue is computed for STARTER only - '32% of revenue (EUR792/customer/month vs EUR2,462 ARPU)', both exact. Enterprise's support is discussed only as a margin movement and Pro's only as a month-on-month change; the raw ticket counts by plan never appear, so the conclusion that Enterprise is not the worst despite raising 35 of 74 October tickets is never reached. That is the rubric's 'computes a rate for one plan only'
support cost as a share of plan revenue is computed for STARTER only - '32% of revenue (EUR792/customer/month vs EUR2,462 ARPU)', both exact. Enterprise's support is discussed only as a margin movement and Pro's only as a month-on-month change; the raw ticket counts by plan never appear, so the conclusion that Enterprise is not the worst despite raising 35 of 74 October tickets is never reached. That is the rubric's 'computes a rate for one plan only'
causal_boundary−1 pts forgone
causal_boundary
PARTIALPARTIALPASS
all three Soylent movements are present and correct - '-27%' usage, 'tickets went 3 to 8', and the EUR2.5k downgrade - and no mechanism is asserted or implied. What the PASS band also requires is the explicit statement that the data does not establish what caused what, and the response never makes it for Soylent; the epistemic caveat it does give is about channel attribution
all three Soylent movements are present and correct - '-27%' usage, 'tickets went 3 to 8', and the EUR2.5k downgrade - and no mechanism is asserted or implied. What the PASS band also requires is the explicit statement that the data does not establish what caused what, and the response never makes it for Soylent; the epistemic caveat it does give is about channel attribution
headline_vs_recurring_growth−1 pts forgone
States both the headline and a normalised figure
PASSPARTIALPASS
the normalised figure is right and its treatment is stated - EUR746.1k to EUR757.7k, +1.6% - and the gap is correctly attributed to the annual invoice and the duplicate. But the headline it contrasts against is the DEDUPLICATED +9.6%; the raw invoice sum's ~+12.0% is never computed. EUR892.9k appears once, in the data-issues section, as a total rather than as a growth rate. That is the rubric's 'gives only one of the two figures'
the normalised figure is right and its treatment is stated - EUR746.1k to EUR757.7k, +1.6% - and the gap is correctly attributed to the annual invoice and the duplicate. But the headline it contrasts against is the DEDUPLICATED +9.6%; the raw invoice sum's ~+12.0% is never computed. EUR892.9k appears once, in the data-issues section, as a total rather than as a growth rate. That is the rubric's 'gives only one of the two figures'
segment_retention_quantified−0.8 pts forgone
Quantifies retention on a stated basis
PASSPARTIALPASS
SMB is correctly identified as worst and Enterprise as zero, but the middle is never quantified: no mid-market churn rate appears anywhere, so the three populations cannot be ranked from the figures given. The two framings are also mixed - '44% of the starter base' is plan, '4 of 12 September SMB logos' is segment - and that segment denominator is wrong (13 SMB accounts predate October, not 12). Neither error changes which population is worst, which is the PARTIAL band
SMB is correctly identified as worst and Enterprise as zero, but the middle is never quantified: no mid-market churn rate appears anywhere, so the three populations cannot be ranked from the figures given. The two framings are also mixed - '44% of the starter base' is plan, '4 of 12 September SMB logos' is segment - and that segment denominator is wrong (13 SMB accounts predate October, not 12). Neither error changes which population is worst, which is the PARTIAL band
absolute_vs_scalable_economics−0.8 pts forgone
absolute_vs_scalable_economics
PASSPARTIALPASS
the contrast is drawn correctly - '87% of revenue, 79% of gross profit - but margin fell 2.7pp' against 'Pro is the margin leader (62.7%)' - and then the product recommendation leans on the absolute figure anyway: 'Plan focus (Product): Enterprise plan ... It carries 79% of gross profit'. That is the PARTIAL band's second clause word for word
the contrast is drawn correctly - '87% of revenue, 79% of gross profit - but margin fell 2.7pp' against 'Pro is the margin leader (62.7%)' - and then the product recommendation leans on the absolute figure anyway: 'Plan focus (Product): Enterprise plan ... It carries 79% of gross profit'. That is the PARTIAL band's second clause word for word
product_focus_supported−0.8 pts forgone
product_focus_supported
PASSPARTIALPASS
the Enterprise plan is chosen over Pro. The other-plan route is open in the rubric, but only when argued on cost-adjusted economics and at least two independent signals AND not resting on absolute revenue. Margin erosion and the open high-severity tickets are real signals; the case is nevertheless led by 'it carries 79% of gross profit' and closed by 'one enterprise churn (stark = EUR1.14M) would erase more ARR than a full quarter of new logos', both absolute-size arguments. Pro is named only as a secondary track
the Enterprise plan is chosen over Pro. The other-plan route is open in the rubric, but only when argued on cost-adjusted economics and at least two independent signals AND not resting on absolute revenue. Margin erosion and the open high-severity tickets are real signals; the case is nevertheless led by 'it carries 79% of gross profit' and closed by 'one enterprise churn (stark = EUR1.14M) would erase more ARR than a full quarter of new logos', both absolute-size arguments. Pro is named only as a secondary track
segment_recommendation−0.5 pts forgone
Does not recommend cutting the segment
PASSPARTIALPASS
SMB is not cut and the missing input is named ('Segment P&L cannot be built: costs are attributed by plan, and plans cut across segments'), but the recommendation is a reduction with conditions attached: 'freeze incremental SMB acquisition spend ... set a January checkpoint - cut only if SMB net revenue retention hasn't improved'
SMB is not cut and the missing input is named ('Segment P&L cannot be built: costs are attributed by plan, and plans cut across segments'), but the recommendation is a reduction with conditions attached: 'freeze incremental SMB acquisition spend ... set a January checkpoint - cut only if SMB net revenue retention hasn't improved'
What it got wrong
No penalties and no false actions in any run.
5
Claude Opus 5closed
90.8
9.5
–
–
46
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
90.0
1.0
108.3s
128.6s
Individual runs
Run 189.5
Run 290.0
Run 390.5
Speed
TTFA
108.3s
Total latency
128.6s
Median output tokens
30,057
Dimension breakdown
20/20
15/15
19/19
12.5/19
18/18
5.5/9
What it missed
absolute_vs_scalable_economics−2.5 pts forgone
absolute_vs_scalable_economics
PARTIALPARTIALPARTIAL
The response correctly distinguishes Enterprise absolute profit from Pro margin, but then leans on Enterprise's absolute pool when choosing product focus.
The contrast is calculated but not kept binding in the final decision.
product_focus_supported−2.5 pts forgone
product_focus_supported
PARTIALPARTIALPARTIAL
The response chooses Enterprise first and Pro second; the choice leans materially on Enterprise's absolute profit pool.
It does not cleanly select the scalable-economics signal the rubric is designed to test.
segment_recommendation−2 pts forgone
Does not recommend cutting the segment
FAILPARTIALPARTIAL
The response says 'The data supports reducing new acquisition investment in SMB/Starter.'
The pack lacks segment CAC/LTV/retention economics sufficient to recommend cutting or defunding SMB.
plan_trajectory_evidence−1.5 pts forgone
Reads the customers, not the plan average
PARTIALPARTIALPARTIAL
Usage/adoption trajectory is discussed, but the forward-looking plan comparison is not fully quantified across the plans.
Trajectory evidence is useful but incomplete.
uncertainty_discipline−1 pts forgone
uncertainty_discipline
PARTIALPARTIALPARTIAL
The response names many real data gaps, but its segment recommendation is stronger than those stated gaps permit.
Uncertainty is identified but not fully preserved in the decision.
product_recommendation−0.5 pts forgone
product_recommendation
PASSFAILPARTIAL
The recommendations section makes Enterprise the primary engineering focus despite the analysis identifying Pro as the strongest scalable economics.
The final product choice conflicts with the strongest signal in its own analysis.
What it got wrong
No penalties and no false actions in any run.
7
Fable 5.1closed
89.3
8.5
–
–
58
DeepSeek V4 Flashdeepseek-v4-flash
86.0
10.5
164.6s
174.4s
Individual runs
Run 191.0
Run 280.5
Run 386.5
Speed
TTFA
164.6s
Total latency
174.4s
Median output tokens
29,942
Dimension breakdown
19/20
12.5/15
16.5/19
18/19
16.3/18
6/9
What it missed
segment_retention_quantified−2.5 pts forgone
Quantifies retention on a stated basis
PARTIALPARTIALPARTIAL
the active-MRR table - Enterprise EUR 565.0k flat, Mid-market EUR 142.0k -> EUR 127.5k (-10.2%), SMB EUR 39.1k -> EUR 39.9k (+2.0%) - plus churned ARR of EUR 303.6k itemised by account
Quantifies segment movement on a stated basis and ranks the three, but the basis is net revenue movement rather than retention: SMB's four churns are offset by its five new logos, so no churn or retention rate is given for any segment.
the channel table gives new-logo and expansion counts but only total attributed revenue: paid search EUR 274,800, partner EUR 166,800, content EUR 109,200, outbound EUR 80,400
Having identified the contamination it never removes it, so no new-logo ARR is computed for any channel and the ranking never reverses - which is how it ends up naming paid search as the best candidate for a test.
segment_recommendation−2 pts forgone
Does not recommend cutting the segment
PARTIALFAILPARTIAL
'Reduce investment in the SMB segment, specifically the Starter plan. The data supports this: 4 of 13 SMB customers churned in October ... reduce spend on Starter acquisition until unit economics and retention improve.'
Recommends the reduction the pack rules out, on a churn rate whose denominator is wrong, and asserts that the data supports it. Its own analysis shows SMB net MRR positive and the segment producing all five of October's new logos.
headline_vs_recurring_growth−1 pts forgone
States both the headline and a normalised figure
PARTIALPASSPASS
'Reported October plan MRR is EUR 757.7k vs EUR 746.1k in September (+1.6%), but that includes customers who churned during October. Active MRR at end-October is EUR 732.4k, a -1.8% decline.'
Two bases, each labelled with exactly what it contains, and the gap between them shown. The raw invoiced comparison the headline would be built on is never computed, so the largest distortion is stepped around rather than measured.
plan_trajectory_evidence−1 pts forgone
Reads the customers, not the plan average
PASSFAILPASS
'No plan is currently exceeding its included request allowance, so overage revenue is not masking these economics.'
True of the three plan averages and false of the customers: Stark 4.29M and Wayne 3.41M exceed the 3M enterprise allowance, and Wonka 1.02M, Aperture 705k and Vandelay 690k exceed Pro's 600k. The claim is then used as an argument, which is the aggregate error the scenario is built around, and it is why the Pro rate-limit signal is never found.
figures_carried_correctly−1 pts forgone
Carries its own figures through correctly
PASSPARTIALPARTIAL
checked against the CSVs: both revenue bases, the segment and plan MRR walks, total requests 31.0M -> 31.5M, requests excluding Acme 26.8M -> 29.2M, tokens 32,134M -> 33,937M, Acme's tokens per request, every attributed-revenue cell, all eight funnel rates, all twelve plan figures, active Pro MRR EUR 51.2k, active Starter EUR 21.7k and Enterprise at 87% of MRR
One figure error: '4 of 13 SMB customers' attaches the Starter plan's population to the SMB segment, which has 18, overstating segment churn as 30.8% against 22.2%. It changes no decision - the recommendation it supports is about the Starter plan, for which 13 is the right denominator. The allowance claim is a wrong inference from three correctly quoted averages, so it is charged under plan_trajectory_evidence rather than here.
acquisition_recommendation−1 pts forgone
acquisition_recommendation
PASSPASSFAIL
'If a controlled test is required, paid search is the best candidate because it drove the most attributed revenue and the only high-value expansions.'
Refuses to increase spend, which is right, and then names the one channel the pack rules out - on the raw attributed revenue whose expansion contamination it had identified two sections earlier. Never having computed new-logo ARR by channel is what leaves it there.
support_normalized−0.7 pts forgone
Compares support load on a rate, not a raw count
PARTIALPASSPASS
'Rate-limit tickets by plan: Pro plan went from 5 in September to 13 in October' and 'Wonka alone had 4 rate-limit tickets in October'
Correct counts and a real within-plan trajectory, but no rate against any denominator and no comparison across plans. It avoids the raw-count trap without doing the normalisation.
What it got wrong
PROHIBITED_ACTION-7 · Run 2
'Reduce investment in the SMB segment, specifically the Starter plan. The data supports this: 4 of 13 SMB customers churned in October ... reduce spend on Starter acquisition until unit economics and retention improve.'
Defunding part of SMB on a logo churn rate is the first action the pack rules out, and this states that the data supports it. The denominator is wrong - 13 is the Starter plan, not the 18-customer SMB segment - and the response's own figures show SMB active MRR rising and SMB producing all five October new logos. It carves out 'do not cut all SMB activity', but leadership is still told to cut.
9
Gemini 3.8 Flash (High)closed
78.5
3.5
–
61.9s
610
MIMO v2.5mimo-v2.5
78.5
25.5
368.6s
483.6s
Individual runs
Run 180.0
Run 265.0
Run 390.5
Speed
TTFA
368.6s
Total latency
483.6s
Median output tokens
24,247
Dimension breakdown
18/20
11.8/15
15.8/19
18/19
13.5/18
6/9
What it missed
segment_retention_quantified−2.5 pts forgone
Quantifies retention on a stated basis
PARTIALPARTIALPARTIAL
'Four of nine SMB starter customers (44%) churned in October' and the segment MRR table giving Enterprise EUR 565,000 / Mid-market EUR 129,500 -> EUR 130,000 / SMB EUR 51,600 -> EUR 37,400
The Starter rate is correct and correctly labelled as a plan, and the segments are ranked on MRR movement. But no churn or retention rate is computed for any segment, and the segment table's Mid-market MRR and every customer count in it are wrong, so the segment framing cannot carry the ranking on its own.
figures_carried_correctly−2.5 pts forgone
Carries its own figures through correctly
FAILFAILPARTIAL
'Total tickets 57 -> 74' when September is 59, which then propagates into the cost-per-ticket row; 'starter support cost per EUR 1 of revenue is EUR 0.43' when it is EUR 0.32; 'Three additional expansion bookings (Stark ... Massive-Dynamic = +EUR 192K total)' names four and omits Cyberdyne, so the total is EUR 216K; 'aggregate Enterprise requests fell 1.4%' when the figure is -0.4%; and 'showed -35% to -37% request declines in September' when those declines are October's
Repeated incorrect figures across the support, acquisition and plan sections, and the EUR 0.43 ratio sits inside the argument for reducing SMB investment. The MRR table, the per-customer economics, the churn ARR walk and the plan tables are all exact.
headline_vs_recurring_growth−2 pts forgone
States both the headline and a normalised figure
PARTIALPARTIALPASS
'Total invoiced EUR 797,100 -> EUR 860,900, +8.0%' against 'October subscription MRR was EUR 757,700, up 1.6% from September'
The structure is right - a raw basis and a normalised one, both labelled, with the gap attributed to the annual invoice and the new logos. But its October invoiced total is EUR 860,900 where the deduplicated figure is EUR 873,400, so the headline reads +8.0% instead of about +12% and falls outside the stated tolerance.
segment_recommendation−2 pts forgone
Does not recommend cutting the segment
FAILFAILPASS
'Reduce investment in SMB - YES, proceed ... Pause or reduce paid acquisition targeting SMB/starter segments ... Evaluate whether the Starter plan should be sunset or restructured'
The bluntest version in the battery: it recommends reducing investment in SMB as a segment, not only the Starter plan, and floats sunsetting. Its own figures show SMB producing all of October's new logos, and no segment-level acquisition economics exist in the pack.
the channel table carries leads, trial rates and October attributed revenue only; no new-logo ARR is computed for any channel
Having seen that paid search's number is mostly expansion, it never recomputes the ranking on new-logo revenue, so the reversal never appears.
customer_id_alias_resolved−1.3 pts forgone
Joins billing to usage through the alias
PASSPARTIALPARTIAL
'Wayne Enterprises has an imminent renewal (November 30, EUR 1,008K ARR) ... Despite healthy usage growth (+10%)' - Wayne's usage is read and reported correctly
Wayne's usage growth is right and requires the alias, but no figure is given for it and no aggregate in the response is reachable only with the join, so the evidence is thinner than run-004's exact enterprise request total.
acquisition_recommendation−1 pts forgone
acquisition_recommendation
PARTIALPARTIALPASS
'Content produced the most new logos (3 in 2 months) with stable conversion rates ... This may be more capital-efficient, but we can't prove it without spend data' with no channel increase recommended
Identifies content on the corrected new-logo basis and gates everything on the missing cost data, but stops short of a recommendation - which the criterion places at PARTIAL.
plan_trajectory_evidence−1 pts forgone
Reads the customers, not the plan average
PASSPARTIALPARTIAL
'Highest plan utilisation (74%) - customers are approaching their included request limits' and 'Utilisation is only 42%' for Starter
Utilisation is quantified for every plan and hedged correctly - 'approaching' rather than 'below' - so it avoids the aggregate FAIL. But it reads only the averages: no customer above its included volume is named anywhere, and no usage growth rate is given.
'Three of its four October conversions are expansions, not new logos - expansions do not expand the customer base' and 'High attributed revenue could simply reflect high spend'
The contamination is identified for Paid Search and the reasoning is right, but it is never quantified: neither the EUR 274,800 attributed total nor the EUR 210,000 of expansion appears, only the October column.
acme_false_positive_avoided−0.7 pts forgone
Does not treat the explained usage drop as a risk
PASSPARTIALPASS
'Acme's migration proves usage can fall materially without revenue falling' and, in the uncertainties, 'Cannot attribute Enterprise cost growth ... Acme's usage fell'
The account is never treated as a risk and revenue stability is noted, but neither the -45% figure nor the batch migration is stated in the analysis itself, so the connection the criterion asks for is left implicit and no second usage metric is offered.
'Paid Search 520 leads / 8.5% -> 900 leads / 6.4%' with 'its lead-to-trial rate dropped from 8.5% to 6.4%, suggesting lead quality may be declining as spend scales'
Rising volume against a falling rate is stated for paid search, but the table drops the trial counts, so the other half - that absolute conversions rose while the rate fell - is not shown.
support_normalized−0.7 pts forgone
Compares support load on a rate, not a raw count
PASSPASSPARTIAL
the churn table's per-account ticket counts and 'Wonka generated 8 support tickets in October'
Support is read at account level throughout and Enterprise is never called the worst on raw volume, but no rate against any denominator is computed for any plan or segment, so the normalisation the criterion asks for is not performed.
What it got wrong
PROHIBITED_ACTION-7 · Run 1
'Decision 1: Reduce investment in one customer segment. Recommendation: Yes - reduce investment in SMB starter customers ... The recommendation is to shift resources away from acquiring new starter-tier customers.'
Defunding SMB acquisition on a churn rate is the first action the pack rules out. This is the most carefully bounded version of it in the battery - it names the funnel consequence, preserves SMB Pro and offers alternatives - but the instruction to shift acquisition spend away still stands, and no segment-level acquisition cost exists in the pack to support it.
PROHIBITED_ACTION-7 · Run 2
'Decision 1: Reduce investment in SMB - YES, proceed ... Pause or reduce paid acquisition targeting SMB/starter segments ... Evaluate whether the Starter plan should be sunset or restructured.'
Cutting SMB investment on a churn rate is the first action the pack rules out, and this states it at the segment level rather than the plan level and adds a sunset option. Its own analysis shows SMB is where every new logo came from, and the pack contains no segment-level acquisition cost against which the cut could be justified.
711
GLM 5.2glm5.2
71.2
4.5
57.1s
103.1s
Individual runs
Run 168.5
Run 272.0
Run 373.0
Speed
TTFA
57.1s
Total latency
103.1s
Median output tokens
8,697
Dimension breakdown
17/20
13.5/15
15/19
14.7/19
13.7/18
4.3/9
What it missed
headline_vs_recurring_growth−3 pts forgone
States both the headline and a normalised figure
PARTIALPARTIALPARTIAL
'September total invoiced ~ EUR 797k (subscriptions EUR 746k + professional services EUR 51k) ... October invoiced is much higher' and 'Underlying monthly subscription revenue ex-Gringotts-annualisation is roughly flat to slightly up'
The normalised side passes as written - the criterion accepts 'roughly flat' - and the treatment is stated. But October's total is never given and the headline growth rate never computed, so only one of the two figures is on the page.
segment_recommendation−3 pts forgone
Does not recommend cutting the segment
FAILFAILFAIL
'Decision 1: Reduce investment in SMB - Supported. Recommendation: Reduce customer-success and support investment in the SMB/Starter segment ... reallocate CS capacity' and, in the executive summary, 'Reduce investment in SMB (specifically the Starter plan).'
Recommends the reduction the pack rules out and marks it as supported. It is the most carefully bounded version glm5.2 produced - acquisition is explicitly preserved, the move is to cheaper support tiers rather than an exit, and a churn threshold is set for a pricing review - but leadership is still told to pull CS investment out of the segment, and the pack contains no segment-level economics against which that could be justified. Both sibling runs were failed for the same action.
no new-logo ARR figure for any channel; the acquisition section is the funnel table plus prose
The ranking is never recomputed on new-logo revenue, so the reversal never appears.
product_focus_supported−2.5 pts forgone
product_focus_supported
PASSPARTIALFAIL
'Recommendation: Focus next quarter's engineering on the Enterprise plan' with the first rationale 'Enterprise generates 87% of revenue and has 0% churn - protecting it is the highest-leverage engineering investment.'
A plan other than Pro is admissible only when the case does not rest on absolute revenue, and here the revenue share is the stated priority and the leading argument. The supporting signals it adds - the Wayne renewal, the batch workload, the open tickets - are real, but they are attached to a choice already made on size. Gemma 4 run-005 was failed for the same construction. Pro is correctly analysed and then demoted to 'secondary focus'.
figures_carried_correctly−2 pts forgone
Carries its own figures through correctly
PARTIALFAILPARTIAL
'GP per Starter customer is EUR 792/year' - it is EUR 792 per month, and the annual figure makes the plan look twelve times worse in the argument for cutting it; 'Wonka Industries generated 7 tickets in October, all rate_limits' when 4 of the 7 are; and the revenue table carries a repeated EUR 700 offset (September subscriptions EUR 746.8k against EUR 746.1k, October ~EUR 868.4k against EUR 867.7k, forward MRR EUR 733.1k against EUR 732.4k)
Repeated incorrect figures across three sections, and the units error sits inside the recommendation it supports. The churn table, the converted-customers table and the plan tables are all exact to the euro, which makes the pattern a lapse rather than a misunderstanding - but the criterion charges repetition.
acquisition_recommendation−1.5 pts forgone
acquisition_recommendation
FAILPARTIALPASS
'increase paid search spend only on the Enterprise-expansion cohort (lookalike/ABM-style)' on the strength of 'EUR 174k of Enterprise expansion ACV in October, which is high-value'
The one channel the pack rules out, recommended on the raw attributed revenue whose expansion contamination the response had itself identified. Partner gets a test budget named on conversion rate rather than on new-logo value, which was never computed.
customer_id_alias_resolved−1.3 pts forgone
Joins billing to usage through the alias
PARTIALPARTIALPASS
account-level billing-to-usage joins are performed for Soylent (EUR 234k ARR against usage -27% and active_users 56 -> 41) and for Acme, but Wayne Enterprises' usage appears nowhere
The FAIL clause applies only when no account-level join is attempted at all, and several are. But nothing here is reachable only with Wayne joined, so there is no evidence either way that the alias was handled.
plan_trajectory_evidence−1 pts forgone
Reads the customers, not the plan average
FAILPASSPASS
'Utilisation is low across all plans: Enterprise avg ~1.89M vs 3M included (63%); Pro ~0.44M vs 0.6M (73%); Starter ~0.08-0.11M vs 0.2M (40-55%). Customers are paying for capacity they don't use - a pricing/packaging opportunity'
Three plan averages, all quoted correctly, and then a claim about the customers that the customers contradict: Stark 4.29M and Wayne 3.41M are over the 3M enterprise allowance, and Wonka 1.02M, Aperture 705k and Vandelay 690k are over Pro's 600k. The claim is used as an argument, which is the FAIL clause as written.
causal_boundary−1 pts forgone
causal_boundary
PARTIALPASSPARTIAL
Soylent's usage decline and ticket cluster are reported together and it is called 'a genuine health risk', but the EUR 22k -> EUR 19.5k contraction is not among them
Two of the three simultaneous movements, with no mechanism asserted. The third - the revenue contraction - is the one that makes the pattern a pattern.
'paid search sourced three large Enterprise expansions (Stark, Umbrella, Globex = EUR 174k ACV). These are high-value, so paid search's value per conversion may still be attractive'
It identifies the expansions and quantifies them at EUR 174k, then argues they make paid search look better rather than that they are not new business. The correction is observed and not applied.
segment_retention_quantified−0.8 pts forgone
Quantifies retention on a stated basis
PASSPARTIALPASS
the churn table placing all five logos with segment, plan, ARR and MRR, plus 'Gross revenue retention (MRR basis) 96.6%' and '4 of ~9 active Starter customers churned'
Company-level GRR and NRR are both computed and internally consistent, and the Starter rate is correctly labelled as a plan. But no rate is given for any segment, so the three populations are never ranked on retention.
absolute_vs_scalable_economics−0.8 pts forgone
absolute_vs_scalable_economics
PASSPASSPARTIAL
The strongest-versus-weakest table holds 'Revenue scale: Enterprise strongest (87%)' against 'Margin %: Enterprise weakest (29.5%), Pro strongest (62.7%)' - then concludes 'Strongest: Enterprise' and opens the engineering rationale with 'Enterprise generates 87% of revenue ... protecting it is the highest-leverage engineering investment.'
Presents both halves and draws the contrast explicitly, then argues the product decision from the absolute revenue figure anyway - the second clause of the PARTIAL band, and the confusion the criterion exists to catch. Gemma 4 run-005 was placed here for the same sequence.
no_churn_overreach−0.7 pts forgone
no_churn_overreach
PASSPARTIALPARTIAL
'Nakatomi churned with an open high-severity api_errors ticket - likely a support-driven churn.'
Offers a cause the pack does not record. 'Likely' marks it as an inference rather than a finding, which is the PARTIAL band, but the pack contains no churn reasons at all.
the funnel table carries lead counts, conversion rates and lead growth, but no absolute conversion counts: 'paid_search 520 / 8.5% / 900 / 6.4% / +73%'
The falling rate against rising volume is stated and correct. The other half - that absolute conversions rose while the rate fell - needs the conversion counts, which this table drops.
product_recommendation−0.2 pts forgone
product_recommendation
PASSPARTIALPASS
'Enterprise, for cost reduction; Pro, for growth enablement'
Actionable and consistent with its own section 4, but two recommendations rather than the single one the criterion asks for.
What it got wrong
PROHIBITED_ACTION-7 · Run 1
'Decision 1 - Reduce investment in a customer segment: Yes, reduce investment in SMB Starter ... Action: stop actively acquiring Starter logos; migrate viable Starter customers up to Pro; tighten support SLA on Starter.'
Defunding part of SMB on a churn rate is the first action the pack rules out, and this is an unhedged instruction to stop acquiring. The supporting figures are computed correctly and the segment's healthy half is preserved, which is why the criterion reads FAIL rather than the analysis being wrong - but the pack records no segment-level acquisition cost against which a spend cut could be justified, and the response never says so.
PROHIBITED_ACTION-7 · Run 2
'Decision 1: Reduce investment in one customer segment - Yes, reduce SMB Starter investment ... Recommendation: Stop actively acquiring Starter customers via paid channels ... redirect acquisition spend and CS capacity toward Pro and Enterprise.'
Defunding part of SMB on churn and plan economics, which is the first action the pack rules out. It is the same recommendation as run-003 and is charged the same way. Part of the case rests on 'GP per Starter customer is EUR 792/year', which is the monthly figure mislabelled - the plan is twelve times more profitable than the sentence justifying the cut says it is.
PROHIBITED_ACTION-7 · Run 3
'Decision 1: Reduce investment in SMB - Supported. Recommendation: Reduce customer-success and support investment in the SMB/Starter segment. Do not stop acquisition, but reallocate CS capacity.' - carried into the executive summary as a decision the data supports.
Pulling CS and support resource out of SMB on a churn rate is the action the pack rules out. It is the most bounded version in glm5.2's three runs - acquisition is preserved, the move is to self-service rather than an exit, and a churn threshold triggers a pricing review instead of a cut - but leadership is still told to withdraw investment from the segment, and no segment-level economics exist in the pack against which that could be justified. qwen3.6 run-005 was penalised for the same reallocation.
812
Qwen 3.6qwen3.6
46.0
2.5
72.6s
102.4s
Individual runs
Run 144.5
Run 247.0
Run 346.5
Speed
TTFA
72.6s
Total latency
102.4s
Median output tokens
8,548
Dimension breakdown
9.7/20
9.3/15
10.5/19
13.7/19
5.5/18
4.3/9
What it missed
headline_vs_recurring_growth−5 pts forgone
States both the headline and a normalised figure
PARTIALFAILFAIL
'Monthly subscription revenue is stable: EUR 659k (Sept) -> EUR 659.5k (Oct) for Enterprise; EUR 63k -> EUR 66.2k for Pro; EUR 24.1k -> EUR 32k for Starter'
Three plan lines, all exact, and no company total on any basis. Neither the headline growth nor a normalised figure refuting it is ever computed, so there is no month-on-month comparison to make.
customer_id_alias_resolved−4 pts forgone
Joins billing to usage through the alias
FAILFAILFAIL
no account-level join between billing and usage is attempted; Wayne appears only in a list of accounts with high-severity tickets, which comes from support.csv alone
The FAIL clause applies: with no billing-to-usage join attempted there is no evidence the alias was handled.
sample_size_arr_context−3.3 pts forgone
sample_size_arr_context
PARTIALFAILFAIL
'support costs that are disproportionately high relative to ARR' and 'low usage relative to included limits' as the context for the 44%
The churn is never sized in ARR, never netted against the new logos, and never compared with the Mid-market loss. The context offered is plan economics rather than the size or value of what was lost.
the channel table reports attributed revenue and revenue per conversion over both months combined - partner EUR 166.8k, paid search EUR 274.8k - with no new-logo column
No new-logo ranking is computed. Partner leads its table on efficiency ratios rather than on the reversal, and paid search's EUR 274.8k is still carried as its revenue.
segment_recommendation−3 pts forgone
Does not recommend cutting the segment
FAILFAILFAIL
'Reduce investment in the SMB/Starters segment: Reallocate CS and marketing resources away from cold outreach to Starter-tier prospects.'
Recommends the reduction on churn and support cost. It softens it - retention effort stays with existing accounts, and a price review is proposed - but resources are moved out of the segment, and the pack has no segment-level economics to support the move.
the funnel table is 'Sept + Oct Combined' - Partner 135 leads / 32 conversions / 23.7%, Paid Search 1,420 / 102 / 7.2%
Combining the months erases the movement entirely. Neither the falling conversion rate nor the rising absolute conversions is reported, and the pack's central funnel observation is never made.
'Paid Search ... drove EUR 216k in expansion bookings in Oct' and 'Paid Search is the most scalable channel and is successfully driving expansion revenue'
The expansion is separated from the new-logo figure, which is the right move. But it is quantified wrongly - EUR 216k is every channel's October expansion, not paid search's EUR 174k - and the conclusion drawn is that paid search is succeeding, not that its headline is not new business.
causal_boundary−2.5 pts forgone
causal_boundary
FAILFAILPARTIAL
Soylent appears once, in the revenue section: 'Soylent's invoice dropped from EUR 22k to EUR 19.5k, reflecting usage or contract adjustments'. Its usage decline and its ticket spike are absent
One of the three simultaneous movements, which is the FAIL bar. The pattern is never assembled - and 'reflecting usage or contract adjustments' offers a mechanism for the one movement it does report.
annual_prepayment_normalized−2 pts forgone
Normalises the annual prepayment
PARTIALPASSPARTIAL
'Gringotts annual contract (EUR 120k) inflates October cash collection. Normalized to MRR (EUR 10k/mo), October MRR = EUR 767,700.'
The invoice is identified and its monthly equivalent stated correctly - and then added to a figure that already contains it. EUR 757,700 is the plans.csv MRR, which counts Gringotts at EUR 10,000 once; adding EUR 10,000 again produces a total that exists nowhere. The normalisation is named and misapplied.
revenue_conclusion−2 pts forgone
revenue_conclusion
PARTIALPARTIALPARTIAL
'Revenue grew modestly (~2.9% MoM on normalized MRR) ... However, gross profit actually contracted as infrastructure and support costs outpaced revenue growth.'
The gross-profit contraction is a real and well-supported finding, but the section's headline conclusion is that revenue grew - which is the reading the scenario exists to refute. 'Growth is flat' appears in the workings and does not reach the summary.
plan_trajectory_evidence−2 pts forgone
Reads the customers, not the plan average
FAILPARTIALPARTIAL
'Usage remains under the 3M included limit, so no overage revenue offsets costs' for Enterprise
The FAIL clause as written: a claim about utilisation that the accounts contradict. Stark used 4.29M and Wayne 3.41M against a 3M allowance, and the claim is then used as an argument for why Enterprise margin is compressed.
support_normalized−2 pts forgone
Compares support load on a rate, not a raw count
PARTIALPARTIALPARTIAL
'Support costs per customer (EUR 790) are high relative to revenue' for Starter
One plan normalised against customers and revenue, with no comparison across plans - and Enterprise's support load is described as 'high but expected for large accounts', which is the right instinct stated without a rate.
cross_table_customer_signal−2 pts forgone
cross_table_customer_signal
PARTIALPARTIALPARTIAL
'Soylent (Mid-market) raised a critical API error ticket in Oct' and the SMB churners read across the churn list and the usage file
Two tables joined for several accounts with correct figures, but no account is built from three or more: Soylent's usage decline and its invoice reduction are both absent.
figures_carried_correctly−2 pts forgone
Carries its own figures through correctly
FAILPARTIALPARTIAL
'October MRR = EUR 767,700' double-counts Gringotts and produces the +2.9% the summary leads with; 'Mid-market ... revenue grew (EUR 63k -> EUR 66.2k)' is the Pro plan's revenue, not the Mid-market segment's; 'SMB (9 -> 13 customers)' and 'Mid-market (8 -> 9)' are plan counts; 'Tickets increased 59 -> 73' when October is 74; and two of the four new-logo ARR averages are wrong
Repeated incorrect figures across four sections, and the first of them carries the executive summary's headline growth rate. The plan economics table is exact throughout, which makes the pattern a failure of labelling and joining rather than of arithmetic.
segment_retention_quantified−1.7 pts forgone
Quantifies retention on a stated basis
PARTIALPASSPARTIAL
'SMB (9 -> 13 customers): 4/9 customers churned in October (44% monthly churn rate)' with 'Enterprise (14 customers): Zero churn' and 'Mid-market (8 -> 9 customers): One churn'
A rate is computed for the worst population and the three groups are ranked correctly. But the labels are plan counts wearing segment names throughout - 9 -> 13 is the Starter plan, 14 is the Enterprise plan, and Mid-market has 10 customers, not 8 - so the framings are mixed without saying so.
absolute_vs_scalable_economics−1.7 pts forgone
absolute_vs_scalable_economics
PARTIALPASSPARTIAL
'Enterprise is the cash engine but margin compressed by EUR 2.7pp' against 'Pro is the most profitable plan (~63% margin) with stable usage and low churn. It deserves growth investment but is small.'
Both halves are present and the contrast is drawn - including that Pro's size is not a reason against it. But it does not say which question each figure answers, and the engineering recommendation then goes to Starter and Enterprise rather than following the economics it just laid out.
product_focus_supported−1.7 pts forgone
product_focus_supported
PARTIALPASSPARTIAL
'Prioritize Starter retention and infra cost optimization for Enterprise' - two plans, and neither is the Pro plan it had just called the most profitable
Names two plans without choosing between them, which is the PARTIAL band. The Starter half is argued from churn and support cost, both cost-adjusted; the Enterprise half from an infra cost rise it quantifies. Neither rests on absolute revenue, so it stays off FAIL.
acquisition_recommendation−1.5 pts forgone
acquisition_recommendation
PARTIALPARTIALPARTIAL
'Increase spend on Partner channel. It delivers 3x the conversion rate and 2.5x the ARR per new logo vs. Paid Search' with 'allocate budget based on conversion rate and ARR quality' until spend data arrives
Partner is the right channel and it is chosen partly on new-logo value. But the increase is not conditioned on the missing cost data - it explicitly proposes allocating budget without it - which is what the criterion asks for.
duplicate_detected−1.3 pts forgone
Finds the duplicated invoice
PASSFAILPASS
no repeated transaction id anywhere; the uncertainties section raises six other issues and not this one
The duplicate is not mentioned.
acme_false_positive_avoided−0.7 pts forgone
Does not treat the explained usage drop as a risk
PARTIALPASSPASS
'Acme's 45% usage drop aligns with the operational note' in the uncertainties section
The migration is named and Acme is kept off every risk list, but no second signal - tokens, revenue or tickets - is offered, and the point is made in the caveats rather than in the customer-health section where the reader would look for it.
product_recommendation−0.2 pts forgone
product_recommendation
PARTIALPASSPASS
'Prioritize Starter retention and infra cost optimization for Enterprise'
Actionable, but two recommendations rather than one, and neither is the plan its own section 4 identifies as the healthiest.
What it got wrong
PROHIBITED_ACTION-7 · Run 1
'Customer Segment: Reduce/redirect investment in SMB. The 44% monthly churn rate and pre-churn usage drop signal a product-market fit or onboarding failure. Shift growth budget to Mid-market and Enterprise, where retention is strong and LTV is higher.'
Defunding SMB on a logo churn rate is the first action the pack rules out. The 44% is the Starter plan's rate presented as the segment's, LTV by segment is not in the pack at all, and the response's own figures show SMB producing every new logo of the period.
PROHIBITED_ACTION-7 · Run 2
'Reduce Investment in SMB (Starter Segment): Halt active growth spend here. Shift to a fully automated, self-serve onboarding flow ... Reallocate budget to Pro and Enterprise.'
Halting growth spend in SMB on a churn rate is the first action the pack rules out, and this is unhedged. No segment-level acquisition cost exists in the pack, and the response's own churn table shows the loss is four small Starter accounts against one Mid-market churn worth more than all of them.
PROHIBITED_ACTION-7 · Run 3
'Reduce investment in the SMB/Starters segment: Reallocate CS and marketing resources away from cold outreach to Starter-tier prospects.' - stated as the first line of the executive summary and repeated as recommendation 1.
Moving marketing and CS resources out of SMB on a churn rate is the action the pack rules out. It is the most carefully bounded of qwen3.6's three versions - retention effort stays with existing accounts and a pricing review is proposed instead of an exit - but leadership is still told to pull resources, and no segment-level economics exist to justify it.
913
Gemma 4gemma4
35.2
19.0
0.8s
15.3s
Individual runs
Run 147.5
Run 228.5
Run 329.5
Speed
TTFA
0.8s
Total latency
15.3s
Median output tokens
1,394
Dimension breakdown
6.7/20
7.2/15
4.7/19
10.8/19
2.5/18
5.7/9
What it missed
headline_vs_recurring_growth−5 pts forgone
States both the headline and a normalised figure
PARTIALFAILFAIL
'Headline revenue is heavily skewed by the Enterprise segment' with no revenue total on any basis
Neither the raw nor the recurring figure is computed, so no comparison exists.
existing_customer_expansion_removed−5 pts forgone
Recomputes the channel ranking on new logos
FAILFAILFAIL
no new-logo ARR figure for any channel
Having identified that expansion inflates paid search, it never removes it, so the ranking never reverses.
attributed revenue is used as-is throughout the acquisition section
Expansion is never separated and the paid-search expansion block is never mentioned.
duplicate_detected−4 pts forgone
Finds the duplicated invoice
FAILFAILFAIL
no repeated transaction id anywhere in the response
The duplicate is not mentioned.
customer_id_alias_resolved−4 pts forgone
Joins billing to usage through the alias
FAILFAILFAIL
Wayne appears only in a list of accounts with high-severity tickets; no account-level usage join is attempted
The alias is neither resolved nor noticed.
support_normalized−4 pts forgone
Compares support load on a rate, not a raw count
FAILFAILFAIL
'Support data shows a high concentration of High severity tickets ... among Enterprise clients (Acme, Stark, Wayne, Umbrella)'
No rate against any denominator, and the engineering recommendation is argued from Enterprise's raw ticket volume - 2.50 per customer against Pro's 3.11.
sample_size_arr_context−3.3 pts forgone
sample_size_arr_context
FAILFAILPARTIAL
'SMB: Low ARR, high churn rate' with no figure
None of the three contexts the criterion asks for is present.
Correct for October and September is absent, so the movement in either direction is invisible.
product_focus_supported−3.3 pts forgone
product_focus_supported
PARTIALPARTIALFAIL
'Direct Engineering focus to the Enterprise Plan ... Protecting this EUR 659k/month revenue is the priority.'
A plan other than Pro is acceptable only when the case does not rest on absolute revenue. Here the revenue figure is the stated priority and the stability argument rests on raw ticket counts.
plan_trajectory_evidence−3 pts forgone
Reads the customers, not the plan average
FAILFAILFAIL
no usage trajectory and no reference to included request volumes
Neither growth nor the distribution against the allowance is examined.
causal_boundary−3 pts forgone
causal_boundary
FAILFAILFAIL
Soylent appears nowhere in the response
None of the three simultaneous movements is reported, so the pattern is never seen.
annual_prepayment_normalized−3 pts forgone
Normalises the annual prepayment
PASSFAILPARTIAL
'Revenue Growth is driven by expansion and new annual contracts, not organic subscription growth' - the Gringotts invoice is never identified
Gestures at annual contracts without finding the EUR 120,000 invoice, its monthly equivalent or its effect on the month.
segment_retention_quantified−2.5 pts forgone
Quantifies retention on a stated basis
PARTIALPARTIALPARTIAL
'SMB: High churn risk. Four SMB customers churned in October alone (Slate-rock, Mega-lo-mart, Dinoco, Buy-n-large)'; Enterprise 'low churn'; Mid-market 'churn (Nakatomi)'
The churned logos are placed in the right segments, but no rate is computed on any basis and the segments are never ranked on retention.
figures_carried_correctly−2.5 pts forgone
Carries its own figures through correctly
FAILFAILPARTIAL
'Partner: EUR 1,110,000 / 135 leads = EUR 8,222/lead' when partner's attributed ARR across both months is EUR 166,800; 'Paid Search: EUR 332,800' when it is EUR 274,800; and five of the six plan-margin figures wrong
Repeated incorrect figures across three different sections, and both the plan conclusion and the channel recommendation stand on them.
acme_false_positive_avoided−2 pts forgone
Does not treat the explained usage drop as a risk
PARTIALPARTIALPARTIAL
'planned shift to batch processing' and, in the uncertainties, 'we cannot confirm if this will lead to a revenue decrease at renewal'
Names the migration and does not treat Acme as a churn risk, but offers no second metric - tokens or revenue - to show the workload moved rather than shrank.
cross_table_customer_signal−2 pts forgone
cross_table_customer_signal
PARTIALPARTIALPARTIAL
Acme read across usage and the operational note, with the renewal implication drawn
One account from two sources. No account is assembled from three or more tables, and Soylent is not mentioned.
acquisition_recommendation−2 pts forgone
acquisition_recommendation
PARTIALPARTIALFAIL
'Increase spend on Paid Search. It is the primary driver for both new logos and high-value expansion.'
The wrong channel, on the raw attribution it had itself flagged as contaminated, and against its own table showing paid search with the worst conversion rate and the steepest decline.
'Decision Support: Increase spend on Paid Search. While Partner has a higher conversion rate, Paid Search is successfully capturing both new logos and massive expansion revenue from existing Enterprise clients.'
Recommends scaling the one channel the pack rules out, on the attributed revenue it had itself flagged as contaminated by expansion, and without the cost data it says elsewhere is missing.
revenue_conclusion−1.3 pts forgone
revenue_conclusion
PASSPARTIALPARTIAL
'Revenue Growth is driven by expansion and new annual contracts, not organic subscription growth'
The right direction, stated first, with nothing quantified behind it and the annual invoice never located.
segment_recommendation−1 pts forgone
Does not recommend cutting the segment
PASSPARTIALPARTIAL
'Do not reduce investment in Mid-market yet ... However, investigate the high support costs/ticket volume in this segment.'
Does not recommend cutting SMB, which keeps it off FAIL, but answers about Mid-market and leaves SMB without a next step or a statement of what the pack cannot settle.
absolute_vs_scalable_economics−0.8 pts forgone
absolute_vs_scalable_economics
PASSPASSPARTIAL
'The Pro plan is the most economically efficient. However, the Enterprise plan is the most expensive to run' followed by 'Engineering focus should be on the Enterprise plan ... Protecting this EUR 659k/month revenue is the priority.'
States the distinction correctly and then argues the product decision from the absolute revenue figure, which is the confusion the criterion exists to catch.
contribution_calculated−0.5 pts forgone
contribution_calculated
PASSPARTIALPASS
'Enterprise EUR 204,700 | Pro EUR 41,500 | Starter EUR 17,100' for October, and EUR 211,200 / EUR 43,500 / EUR 8,300 for September
One of six figures is right - October is EUR 194,700 / EUR 41,500 / EUR 10,300 - but the cost-adjusted calculation is performed for every plan on the right definition, which is more than the FAIL bar of comparing plans on revenue alone.
Margins are given for all three plans and the ordering is right, but two figures are materially wrong: the correct margins are 29.5% / 62.7% / 32.2%, and Starter is out by 21 points in the direction that makes the plan look healthy.
cac_payback_unavailable−0.3 pts forgone
Refuses the metric the pack cannot support
PASSPARTIALPASS
'We cannot calculate true CAC Payback because the marketing spend (cost) per channel was not provided in the data pack. We have used Revenue per Lead as a proxy.' followed by a four-row proxy table
The refusal and the missing input are both right, and then it computes the proxy - which the criterion places at PARTIAL - on figures that are themselves wrong, including EUR 1,110,000 of partner revenue that does not exist.
uncertainty_discipline−0.3 pts forgone
uncertainty_discipline
PASSPARTIALPASS
four items: the causal direction between support and churn, whether Acme's drop is entirely batching, the missing channel spend, and the expansion-booking lag
The four gaps are real and material. But this criterion also requires that the response invent no figures anywhere, and 'Partner: EUR 1,110,000 / 135 leads' is a channel total nearly seven times the EUR 166,800 in the pack.
What it got wrong
PROHIBITED_ACTION-7 · Run 3
'Decision Support: Increase spend on Paid Search.' and, in the recommendations, 'Increase spend on Paid Search. It is the primary driver for both new logos and high-value expansion.'
Materially increasing paid-search spend on raw attributed revenue is the second action the pack rules out, and this response reaches it after correctly noting in its own table footnote that paid search's attributed revenue is inflated by expansion bookings from existing enterprise accounts. Its own figures also show paid search with the worst conversion rate of the four channels and the steepest decline. Naming the option to reject it would have been right; this tells leadership to do it.
Run range: the difference between the model’s best and worst official run. Not a deviation from the mean, so it is never written as ±.
TTFA, time to first answer: the wall-clock delay before the first visible answer token. Reasoning tokens do not count as answer content.
// quality vs speed
The smartest model isn’t always the one you want to wait for
GLM 5.3 Flash tops this scenario at 98.0, after 766.2s before its first answer token. GPT-6 Astra starts answering in — and scores 98.0. Models trade quality for response time in very different ways.
Two independent axes. There is no combined score, and there will not be one: how good the briefing is and how long you wait for it are different questions, and which matters more depends on what you are building.
Monday Score and time to first answer, per model
Model
Monday Score
TTFA median
GLM 5.3 Flash
98.0
766.2s
GPT-6 Astra
98.0
—
Qwen 3.8 Flash
96.5
394.1s
GLM 5.3
91.7
634.2s
DeepSeek V4.1 Flash
90.0
108.3s
DeepSeek V4 Flash
86.0
164.6s
MIMO v2.5
78.5
368.6s
GLM 5.2
71.2
57.1s
Qwen 3.6
46.0
72.6s
Gemma 4
35.2
0.8s
// what stood out
Three things worth saying out loud
Every figure below is one you can check against the leaderboard, the criterion matrix or the published verdicts on the same page.
01
9 / 24
runs recommended the cut the pack rules out
The SMB segment
They did the analysis correctly and then argued against it
Nine of twenty-four runs recommend reducing spend on SMB. Several of them compute, in the same answer, that SMB net MRR is positive and that the segment produced every one of October’s five new logos — and recommend the cut anyway, on a churn rate whose denominator counts the wrong population. The failure is not arithmetic. Each of these runs can read a table; what none of them does is notice that its own numbers have just contradicted its recommendation.
02
7 / 24
runs saw the contamination and left it in
Paid search
Spotting the problem and fixing it are scored separately, and they came apart
The acquisition table mixes new-logo revenue with expansion from customers a channel already had. Paid search leads it, and stops leading the moment expansion is removed. Seven runs identify the contamination in words — “paid search’s number is mostly expansion” — and then never recompute the ranking, so the reversal never appears and the recommendation still names paid search. A benchmark that only asked whether the model noticed would have scored these as passes.
03
62.8
points between first and last — the widest of the five Mondays
The field
The Monday that decides the leaderboard
GLM 5.3 Flash scores 98.0 here and Gemma 4 scores 35.2 — a spread more than three times #005’s, and the worst result of the suite for five of the eight models. It is also where the ranking is settled: five of the six models below the top two record their lowest Monday on this scenario. Give a model tables instead of prose and the field stops looking like a field.
02
The test
What they were given, and what they had to work out.
// the test
Four bases. None of them false
October’s invoices really do add up to €892,900, and that really is 12% more than September. The duplicate really is in the file. Gringotts really did pay €120,000. Nothing in this scenario is a lie — the growth is.
Invoiced, as loadedwhat a dashboard reports
September€797,100
October€892,900
+12.0%
-€19,500INV-2026-10-0209 — Soylent, loaded from both stripe and billing_sync
After removing the duplicate invoice
September€797,100
October€873,400
+9.6%
-€110,000Gringotts prepaid €120,000 for twelve months; €10,000 belongs to October
Normalised: the annual invoice at its monthly valueaccepted
September€797,100
October€763,400
−4.2%the sign changes here
Recurring subscriptions onlyaccepted
September€746,100
October€757,700
+1.6%
-€25,300Five customers were invoiced on 1 October and churned later that month
Active run-rate, once October’s churn drops outrun-rate
September€746,100
October€732,400
−1.8%
Two of these are accepted answers, not one. Normalising the annual invoice while keeping one-off services gives −4.2%; the recurring line alone gives +1.6%. September carried €51,000 of professional services against October’s €5,700, which is why they disagree. Both refute the headline, so the rubric takes either — provided the response says which it computed.
Three more of the same shape
Eight files: a context note and seven CSVs. Every one is internally consistent, and not one of them contains a decision.
The average that is true and useless
Every plan sits below its included request volume on average — 63%, 74%, 42%. Five customers are over it: Stark at 143% of a 3M allowance, Wonka at 170% of 600k. “No plan exceeds its quota” is true of the averages and false of the customers, and a response that argues from it has read the right column of the wrong table.
The channel that leads on the wrong metric
Paid search tops the attributed-revenue table at €274,800. €210,000 of that is expansion on Stark, Umbrella, Globex and Hooli — customers since 2023 and 2024, credited to a campaign that started on 28 September. On new logos alone paid search is third, and partner is first.
The join that silently drops €1M
Two accounts use a different identifier in the usage file than in the billing file — wayne_ent and prestige_ww. Joining on customer_id alone loses Wayne Enterprises, a €1,008,000 contract that renews in four weeks, and no error is raised. The totals simply come out smaller.
The 45% drop that is not a problem
Acme’s requests fall 45%. An operational note says it moved reporting to scheduled batch runs, with an expected 40–50% reduction. Tokens fall only 12% and revenue does not move. The work did not leave; it changed shape. Reporting Acme as a churn risk is the false positive the scenario is built to catch.
03
The method
How the score is built, and how to check it.
// scoring
How Monday Score works
One hundred points across 6 dimensions, then penalties. Quality only: speed and cost are reported beside the score and never inside it.
20151919189
Revenue reasoningDoes it find the growth that is not there?20
Customer healthDoes it size the churn before reacting to it?15
Acquisition reasoningDoes it tell attribution apart from acquisition?19
Plan economicsDoes it know which number the decision depends on?19
Cross-table data qualityDo the joins hold and do the figures survive them?18
Decision qualityDoes it answer the three questions it was asked?9
total100
The judge does not assign a 0–100 score directly
It classifies atomic criteria one at a time, 27 of them in rubric 0.3, and each verdict carries the evidence it was based on. Deterministic scoring code turns those verdicts into points. No model ever sees a running total.
PASSfull points for the criterion
PARTIALhalf points
FAILnothing
Where a response sits between two labels, the rubric’s own tie-break applies: choose the lower one.
Penalties
Applied on top of the dimension score, capped at −25 per run.
-5Material hallucinationA claim that matters to the briefing and is not supported by the material.
-7Prohibited actionRecommending something the current state explicitly rules out.
-7Major data errorA numeric or data mistake that a recommendation or a headline conclusion then rests on.
-5Unsupported causalityAsserting that one thing caused another when the data only shows they moved together.
Every input, raw response, judge verdict and scoring rule for this scenario is published. The hash below is of the assembled prompt each model received, byte for byte.