Nuevo: DeepSeek V4.1 Flash.La arquitectura construida para agentes, y los 890 bytes que cambian la factura.Explorar la pieza →
// mondaybench #004beta
El mismo lunes caótico para 10 modelos abiertos
Ocho ficheros, siete de ellos tablas. Dos meses de ingresos, uso, fugas, captación y soporte de una empresa SaaS, y tres decisiones que esperan a números que no están en ninguna tabla.
Una pregunta: ¿Sabe razonar sobre datos de negocio?
plans.csv 6 filas planingresos, coste de infra y soporte, cuota incluida
acquisition.csv 15 filas customer_idcanal, logo nuevo o expansión, ARR atribuido
funnel.csv 8 filas channelleads y pruebas de pago, por canal y mes
support.csv 133 filas customer_idtickets, severidad, estado
usage_account_idcruza por otra clave: dos cuentas la escriben de otra forma
01
El resultado
Diez modelos, 30 runs, un juez.
// resultados
La clasificación
El Monday Score mide solo calidad: si el briefing reconstruyó bien la situación y eligió el trabajo que importaba. La velocidad va al lado y nunca se mezcla dentro.
Solo modelos abiertosTodos los modelos
monday-001 · 13 models, 3 runs each
#
Modelo
Monday Score
Rango de runs
TTFA mediana
Latencia total
Ver detalle de
11
GLM 5.3 Flashglm5.3-flash
98.0
1.5
766.2s
827.8s
Runs individuales
Run 198.5
Run 298.5
Run 397.0
Velocidad
TTFA
766.2s
Latencia total
827.8s
Mediana de tokens de salida
36,248
Desglose por dimensión
20/20
15/15
19/19
19/19
16.5/18
8.5/9
Qué se dejó
figures_carried_correctly−1.5 pts perdidos
Arrastra sus propias cifras sin corromperlas
PARTIALPARTIALPARTIAL
'wonka ... 7 tickets in October - 5 of them rate_limits' when 4 of the 7 are rate_limits; and 'wayne renews in ~5 weeks' when 2026-11-30 is four weeks after the review date
Two small slips in supporting detail. Everything else checkable is exact, including the EUR 262.9k -> EUR 246.5k profit walk, the +5.8% cost growth, the MRR bridge, the EUR 216k expansion bookings itemised, the EUR 3.23M of ARR with open high-severity tickets, all eight funnel cells, the full new-logo table, all twelve plan figures, contribution per customer for all three plans, and the mid-market -EUR 42k walk.
segment_recommendation−0.5 pts perdidos
No recomienda recortar el segmento
PASSPASSPARTIAL
'Yes - reduce acquisition investment in SMB, redirect rather than cut ... pause incremental paid SMB acquisition, redirect budget to SMB retention/onboarding ... Revisit once CAC-by-segment exists.'
A reduction hedged with conditions, which is the PARTIAL band precisely. It names the missing input the PASS bar asks for and the budget stays inside SMB rather than leaving it, but leadership is still told to pause acquisition spend.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
2
GPT-6 Astragpt-6-astracerrado
98.0
0.0
—
—
Runs individuales
Run 198.0
Run 298.0
Run 398.0
Velocidad
TTFA
—
Latencia total
—
Mediana de tokens de salida
—
Desglose por dimensión
20/20
15/15
19/19
19/19
16/18
9/9
Qué se dejó
support_normalized−2 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PARTIALPARTIALPARTIAL
PARTIAL: Pro's ticket movement is quantified (19 -> 28, rate limits 5 -> 13) and support cost sits per plan in the contribution table, but no rate is computed on any denominator and Enterprise's raw share of tickets is never addressed
PARTIAL: Pro's ticket movement is quantified (19 -> 28, rate limits 5 -> 13) and support cost sits per plan in the contribution table, but no rate is computed on any denominator and Enterprise's raw share of tickets is never addressed
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
23
Qwen 3.8 Flashqwen3.8-flash
96.5
3.0
394.1s
468.0s
Runs individuales
Run 196.5
Run 298.0
Run 395.0
Velocidad
TTFA
394.1s
Latencia total
468.0s
Mediana de tokens de salida
52,295
Desglose por dimensión
20/20
15/15
19/19
19/19
15/18
8.5/9
Qué se dejó
support_normalized−2 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PARTIALPARTIALPARTIAL
'October support tickets by plan: Enterprise 34 -> 35, Pro 19 -> 28, Starter 7 -> 11' and 'September Enterprise-segment tickets: 26, October: 25'
It avoids the trap - Enterprise is called stable and Pro gets the focus - and it reads the segment view as well as the plan view, which is more care than most. But no rate against any denominator is computed, so the normalisation the criterion asks for is not performed.
figures_carried_correctly−1 pts perdidos
Arrastra sus propias cifras sin corromperlas
PARTIALPASSPARTIAL
'Enterprise 34 September tickets' where the enterprise plan has 33, which also makes its own column sum to 60 against a September total of 59
One miscount. Everything else checkable is exact, and the response carries more arithmetic than any other in the battery: the ARR bridge reconciling to EUR 8,788,800 to the euro, four revenue bases, the segment ARR walk including the reconstructed EUR 1,704,000, both channel tables with totals, the full two-month plan economics with cost subtotals, the utilisation table computed from customer counts, all five over-allowance accounts with overage volumes, the Soylent four-metric table, and the 13 October rate-limit tickets itemised per account - the only response anywhere to get that count right.
segment_recommendation−0.5 pts perdidos
No recomienda recortar el segmento
PASSPASSPARTIAL
'the data do not support a full cut without CAC and segment cost data' and 'Reallocate SMB investment toward onboarding, usage activation, and retention measurement rather than cutting the segment entirely' - but also 'For SMB, reduce or pause low-quality acquisition spend, especially starter acquisition, until activation and retention improve.'
A reduction hedged with conditions, which is the PARTIAL band. It names the missing inputs the PASS bar asks for and keeps the budget inside SMB, but leadership is still told to pause acquisition spend rather than to investigate first.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
34
GLM 5.3glm5.3
91.7
17.0
634.2s
673.2s
Runs individuales
Run 195.0
Run 281.5
Run 398.5
Velocidad
TTFA
634.2s
Latencia total
673.2s
Mediana de tokens de salida
46,592
Desglose por dimensión
19/20
14.2/15
19/19
17.3/19
13.7/18
8.5/9
Qué se dejó
figures_carried_correctly−2 pts perdidos
Arrastra sus propias cifras sin corromperlas
PARTIALFAILPARTIAL
five incorrect figures across different parts of the analysis, one of them a headline that contradicts its own components. The ARR bridge reads 'EUR8,953k to EUR8,789k = -EUR164k (-1.8%)' while the components listed beneath it (+246, +133, +36, -30, -304) sum to +EUR81k; the closing figure is right, the opening figure is unreachable, and the stated direction is wrong - ARR on 1 September was about EUR8,713k, so it rose. Also: Acme is put at '84%' of its 3M included requests when 2.31M is 77%; '4 of 12 September SMB logos' where there are 13; rate-limit tickets 'tripled (4 to 13)' where September is 5; and 'infra per token ~EUR11.7/M' which divides the enterprise PLAN's infra cost by the enterprise SEGMENT's tokens, the plan-consistent figure being about EUR9.7/M - a wrong secondary rate. Repeated incorrect figures across different parts, plus one materially wrong: the FAIL band
five incorrect figures across different parts of the analysis, one of them a headline that contradicts its own components. The ARR bridge reads 'EUR8,953k to EUR8,789k = -EUR164k (-1.8%)' while the components listed beneath it (+246, +133, +36, -30, -304) sum to +EUR81k; the closing figure is right, the opening figure is unreachable, and the stated direction is wrong - ARR on 1 September was about EUR8,713k, so it rose. Also: Acme is put at '84%' of its 3M included requests when 2.31M is 77%; '4 of 12 September SMB logos' where there are 13; rate-limit tickets 'tripled (4 to 13)' where September is 5; and 'infra per token ~EUR11.7/M' which divides the enterprise PLAN's infra cost by the enterprise SEGMENT's tokens, the plan-consistent figure being about EUR9.7/M - a wrong secondary rate. Repeated incorrect figures across different parts, plus one materially wrong: the FAIL band
support_normalized−1.3 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PARTIALPARTIALPASS
support cost as a share of plan revenue is computed for STARTER only - '32% of revenue (EUR792/customer/month vs EUR2,462 ARPU)', both exact. Enterprise's support is discussed only as a margin movement and Pro's only as a month-on-month change; the raw ticket counts by plan never appear, so the conclusion that Enterprise is not the worst despite raising 35 of 74 October tickets is never reached. That is the rubric's 'computes a rate for one plan only'
support cost as a share of plan revenue is computed for STARTER only - '32% of revenue (EUR792/customer/month vs EUR2,462 ARPU)', both exact. Enterprise's support is discussed only as a margin movement and Pro's only as a month-on-month change; the raw ticket counts by plan never appear, so the conclusion that Enterprise is not the worst despite raising 35 of 74 October tickets is never reached. That is the rubric's 'computes a rate for one plan only'
causal_boundary−1 pts perdidos
causal_boundary
PARTIALPARTIALPASS
all three Soylent movements are present and correct - '-27%' usage, 'tickets went 3 to 8', and the EUR2.5k downgrade - and no mechanism is asserted or implied. What the PASS band also requires is the explicit statement that the data does not establish what caused what, and the response never makes it for Soylent; the epistemic caveat it does give is about channel attribution
all three Soylent movements are present and correct - '-27%' usage, 'tickets went 3 to 8', and the EUR2.5k downgrade - and no mechanism is asserted or implied. What the PASS band also requires is the explicit statement that the data does not establish what caused what, and the response never makes it for Soylent; the epistemic caveat it does give is about channel attribution
headline_vs_recurring_growth−1 pts perdidos
Da el titular y una cifra normalizada
PASSPARTIALPASS
the normalised figure is right and its treatment is stated - EUR746.1k to EUR757.7k, +1.6% - and the gap is correctly attributed to the annual invoice and the duplicate. But the headline it contrasts against is the DEDUPLICATED +9.6%; the raw invoice sum's ~+12.0% is never computed. EUR892.9k appears once, in the data-issues section, as a total rather than as a growth rate. That is the rubric's 'gives only one of the two figures'
the normalised figure is right and its treatment is stated - EUR746.1k to EUR757.7k, +1.6% - and the gap is correctly attributed to the annual invoice and the duplicate. But the headline it contrasts against is the DEDUPLICATED +9.6%; the raw invoice sum's ~+12.0% is never computed. EUR892.9k appears once, in the data-issues section, as a total rather than as a growth rate. That is the rubric's 'gives only one of the two figures'
segment_retention_quantified−0.8 pts perdidos
Cuantifica la retención sobre una base declarada
PASSPARTIALPASS
SMB is correctly identified as worst and Enterprise as zero, but the middle is never quantified: no mid-market churn rate appears anywhere, so the three populations cannot be ranked from the figures given. The two framings are also mixed - '44% of the starter base' is plan, '4 of 12 September SMB logos' is segment - and that segment denominator is wrong (13 SMB accounts predate October, not 12). Neither error changes which population is worst, which is the PARTIAL band
SMB is correctly identified as worst and Enterprise as zero, but the middle is never quantified: no mid-market churn rate appears anywhere, so the three populations cannot be ranked from the figures given. The two framings are also mixed - '44% of the starter base' is plan, '4 of 12 September SMB logos' is segment - and that segment denominator is wrong (13 SMB accounts predate October, not 12). Neither error changes which population is worst, which is the PARTIAL band
absolute_vs_scalable_economics−0.8 pts perdidos
absolute_vs_scalable_economics
PASSPARTIALPASS
the contrast is drawn correctly - '87% of revenue, 79% of gross profit - but margin fell 2.7pp' against 'Pro is the margin leader (62.7%)' - and then the product recommendation leans on the absolute figure anyway: 'Plan focus (Product): Enterprise plan ... It carries 79% of gross profit'. That is the PARTIAL band's second clause word for word
the contrast is drawn correctly - '87% of revenue, 79% of gross profit - but margin fell 2.7pp' against 'Pro is the margin leader (62.7%)' - and then the product recommendation leans on the absolute figure anyway: 'Plan focus (Product): Enterprise plan ... It carries 79% of gross profit'. That is the PARTIAL band's second clause word for word
product_focus_supported−0.8 pts perdidos
product_focus_supported
PASSPARTIALPASS
the Enterprise plan is chosen over Pro. The other-plan route is open in the rubric, but only when argued on cost-adjusted economics and at least two independent signals AND not resting on absolute revenue. Margin erosion and the open high-severity tickets are real signals; the case is nevertheless led by 'it carries 79% of gross profit' and closed by 'one enterprise churn (stark = EUR1.14M) would erase more ARR than a full quarter of new logos', both absolute-size arguments. Pro is named only as a secondary track
the Enterprise plan is chosen over Pro. The other-plan route is open in the rubric, but only when argued on cost-adjusted economics and at least two independent signals AND not resting on absolute revenue. Margin erosion and the open high-severity tickets are real signals; the case is nevertheless led by 'it carries 79% of gross profit' and closed by 'one enterprise churn (stark = EUR1.14M) would erase more ARR than a full quarter of new logos', both absolute-size arguments. Pro is named only as a secondary track
segment_recommendation−0.5 pts perdidos
No recomienda recortar el segmento
PASSPARTIALPASS
SMB is not cut and the missing input is named ('Segment P&L cannot be built: costs are attributed by plan, and plans cut across segments'), but the recommendation is a reduction with conditions attached: 'freeze incremental SMB acquisition spend ... set a January checkpoint - cut only if SMB net revenue retention hasn't improved'
SMB is not cut and the missing input is named ('Segment P&L cannot be built: costs are attributed by plan, and plans cut across segments'), but the recommendation is a reduction with conditions attached: 'freeze incremental SMB acquisition spend ... set a January checkpoint - cut only if SMB net revenue retention hasn't improved'
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
5
Claude Opus 5cerrado
90.8
9.5
–
–
46
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
90.0
1.0
108.3s
128.6s
Runs individuales
Run 189.5
Run 290.0
Run 390.5
Velocidad
TTFA
108.3s
Latencia total
128.6s
Mediana de tokens de salida
30,057
Desglose por dimensión
20/20
15/15
19/19
12.5/19
18/18
5.5/9
Qué se dejó
absolute_vs_scalable_economics−2.5 pts perdidos
absolute_vs_scalable_economics
PARTIALPARTIALPARTIAL
The response correctly distinguishes Enterprise absolute profit from Pro margin, but then leans on Enterprise's absolute pool when choosing product focus.
The contrast is calculated but not kept binding in the final decision.
product_focus_supported−2.5 pts perdidos
product_focus_supported
PARTIALPARTIALPARTIAL
The response chooses Enterprise first and Pro second; the choice leans materially on Enterprise's absolute profit pool.
It does not cleanly select the scalable-economics signal the rubric is designed to test.
segment_recommendation−2 pts perdidos
No recomienda recortar el segmento
FAILPARTIALPARTIAL
The response says 'The data supports reducing new acquisition investment in SMB/Starter.'
The pack lacks segment CAC/LTV/retention economics sufficient to recommend cutting or defunding SMB.
plan_trajectory_evidence−1.5 pts perdidos
Lee a los clientes, no la media del plan
PARTIALPARTIALPARTIAL
Usage/adoption trajectory is discussed, but the forward-looking plan comparison is not fully quantified across the plans.
Trajectory evidence is useful but incomplete.
uncertainty_discipline−1 pts perdidos
uncertainty_discipline
PARTIALPARTIALPARTIAL
The response names many real data gaps, but its segment recommendation is stronger than those stated gaps permit.
Uncertainty is identified but not fully preserved in the decision.
product_recommendation−0.5 pts perdidos
product_recommendation
PASSFAILPARTIAL
The recommendations section makes Enterprise the primary engineering focus despite the analysis identifying Pro as the strongest scalable economics.
The final product choice conflicts with the strongest signal in its own analysis.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
7
Fable 5.1cerrado
89.3
8.5
–
–
58
DeepSeek V4 Flashdeepseek-v4-flash
86.0
10.5
164.6s
174.4s
Runs individuales
Run 191.0
Run 280.5
Run 386.5
Velocidad
TTFA
164.6s
Latencia total
174.4s
Mediana de tokens de salida
29,942
Desglose por dimensión
19/20
12.5/15
16.5/19
18/19
16.3/18
6/9
Qué se dejó
segment_retention_quantified−2.5 pts perdidos
Cuantifica la retención sobre una base declarada
PARTIALPARTIALPARTIAL
the active-MRR table - Enterprise EUR 565.0k flat, Mid-market EUR 142.0k -> EUR 127.5k (-10.2%), SMB EUR 39.1k -> EUR 39.9k (+2.0%) - plus churned ARR of EUR 303.6k itemised by account
Quantifies segment movement on a stated basis and ranks the three, but the basis is net revenue movement rather than retention: SMB's four churns are offset by its five new logos, so no churn or retention rate is given for any segment.
Rehace el ranking de canales solo con logos nuevos
PASSPARTIALFAIL
the channel table gives new-logo and expansion counts but only total attributed revenue: paid search EUR 274,800, partner EUR 166,800, content EUR 109,200, outbound EUR 80,400
Having identified the contamination it never removes it, so no new-logo ARR is computed for any channel and the ranking never reverses - which is how it ends up naming paid search as the best candidate for a test.
segment_recommendation−2 pts perdidos
No recomienda recortar el segmento
PARTIALFAILPARTIAL
'Reduce investment in the SMB segment, specifically the Starter plan. The data supports this: 4 of 13 SMB customers churned in October ... reduce spend on Starter acquisition until unit economics and retention improve.'
Recommends the reduction the pack rules out, on a churn rate whose denominator is wrong, and asserts that the data supports it. Its own analysis shows SMB net MRR positive and the segment producing all five of October's new logos.
headline_vs_recurring_growth−1 pts perdidos
Da el titular y una cifra normalizada
PARTIALPASSPASS
'Reported October plan MRR is EUR 757.7k vs EUR 746.1k in September (+1.6%), but that includes customers who churned during October. Active MRR at end-October is EUR 732.4k, a -1.8% decline.'
Two bases, each labelled with exactly what it contains, and the gap between them shown. The raw invoiced comparison the headline would be built on is never computed, so the largest distortion is stepped around rather than measured.
plan_trajectory_evidence−1 pts perdidos
Lee a los clientes, no la media del plan
PASSFAILPASS
'No plan is currently exceeding its included request allowance, so overage revenue is not masking these economics.'
True of the three plan averages and false of the customers: Stark 4.29M and Wayne 3.41M exceed the 3M enterprise allowance, and Wonka 1.02M, Aperture 705k and Vandelay 690k exceed Pro's 600k. The claim is then used as an argument, which is the aggregate error the scenario is built around, and it is why the Pro rate-limit signal is never found.
figures_carried_correctly−1 pts perdidos
Arrastra sus propias cifras sin corromperlas
PASSPARTIALPARTIAL
checked against the CSVs: both revenue bases, the segment and plan MRR walks, total requests 31.0M -> 31.5M, requests excluding Acme 26.8M -> 29.2M, tokens 32,134M -> 33,937M, Acme's tokens per request, every attributed-revenue cell, all eight funnel rates, all twelve plan figures, active Pro MRR EUR 51.2k, active Starter EUR 21.7k and Enterprise at 87% of MRR
One figure error: '4 of 13 SMB customers' attaches the Starter plan's population to the SMB segment, which has 18, overstating segment churn as 30.8% against 22.2%. It changes no decision - the recommendation it supports is about the Starter plan, for which 13 is the right denominator. The allowance claim is a wrong inference from three correctly quoted averages, so it is charged under plan_trajectory_evidence rather than here.
acquisition_recommendation−1 pts perdidos
acquisition_recommendation
PASSPASSFAIL
'If a controlled test is required, paid search is the best candidate because it drove the most attributed revenue and the only high-value expansions.'
Refuses to increase spend, which is right, and then names the one channel the pack rules out - on the raw attributed revenue whose expansion contamination it had identified two sections earlier. Never having computed new-logo ARR by channel is what leaves it there.
support_normalized−0.7 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PARTIALPASSPASS
'Rate-limit tickets by plan: Pro plan went from 5 in September to 13 in October' and 'Wonka alone had 4 rate-limit tickets in October'
Correct counts and a real within-plan trajectory, but no rate against any denominator and no comparison across plans. It avoids the raw-count trap without doing the normalisation.
En qué se equivocó
PROHIBITED_ACTION-7 · Run 2
'Reduce investment in the SMB segment, specifically the Starter plan. The data supports this: 4 of 13 SMB customers churned in October ... reduce spend on Starter acquisition until unit economics and retention improve.'
Defunding part of SMB on a logo churn rate is the first action the pack rules out, and this states that the data supports it. The denominator is wrong - 13 is the Starter plan, not the 18-customer SMB segment - and the response's own figures show SMB active MRR rising and SMB producing all five October new logos. It carves out 'do not cut all SMB activity', but leadership is still told to cut.
9
Gemini 3.8 Flash (High)cerrado
78.5
3.5
–
61.9s
610
MIMO v2.5mimo-v2.5
78.5
25.5
368.6s
483.6s
Runs individuales
Run 180.0
Run 265.0
Run 390.5
Velocidad
TTFA
368.6s
Latencia total
483.6s
Mediana de tokens de salida
24,247
Desglose por dimensión
18/20
11.8/15
15.8/19
18/19
13.5/18
6/9
Qué se dejó
segment_retention_quantified−2.5 pts perdidos
Cuantifica la retención sobre una base declarada
PARTIALPARTIALPARTIAL
'Four of nine SMB starter customers (44%) churned in October' and the segment MRR table giving Enterprise EUR 565,000 / Mid-market EUR 129,500 -> EUR 130,000 / SMB EUR 51,600 -> EUR 37,400
The Starter rate is correct and correctly labelled as a plan, and the segments are ranked on MRR movement. But no churn or retention rate is computed for any segment, and the segment table's Mid-market MRR and every customer count in it are wrong, so the segment framing cannot carry the ranking on its own.
figures_carried_correctly−2.5 pts perdidos
Arrastra sus propias cifras sin corromperlas
FAILFAILPARTIAL
'Total tickets 57 -> 74' when September is 59, which then propagates into the cost-per-ticket row; 'starter support cost per EUR 1 of revenue is EUR 0.43' when it is EUR 0.32; 'Three additional expansion bookings (Stark ... Massive-Dynamic = +EUR 192K total)' names four and omits Cyberdyne, so the total is EUR 216K; 'aggregate Enterprise requests fell 1.4%' when the figure is -0.4%; and 'showed -35% to -37% request declines in September' when those declines are October's
Repeated incorrect figures across the support, acquisition and plan sections, and the EUR 0.43 ratio sits inside the argument for reducing SMB investment. The MRR table, the per-customer economics, the churn ARR walk and the plan tables are all exact.
headline_vs_recurring_growth−2 pts perdidos
Da el titular y una cifra normalizada
PARTIALPARTIALPASS
'Total invoiced EUR 797,100 -> EUR 860,900, +8.0%' against 'October subscription MRR was EUR 757,700, up 1.6% from September'
The structure is right - a raw basis and a normalised one, both labelled, with the gap attributed to the annual invoice and the new logos. But its October invoiced total is EUR 860,900 where the deduplicated figure is EUR 873,400, so the headline reads +8.0% instead of about +12% and falls outside the stated tolerance.
segment_recommendation−2 pts perdidos
No recomienda recortar el segmento
FAILFAILPASS
'Reduce investment in SMB - YES, proceed ... Pause or reduce paid acquisition targeting SMB/starter segments ... Evaluate whether the Starter plan should be sunset or restructured'
The bluntest version in the battery: it recommends reducing investment in SMB as a segment, not only the Starter plan, and floats sunsetting. Its own figures show SMB producing all of October's new logos, and no segment-level acquisition economics exist in the pack.
Rehace el ranking de canales solo con logos nuevos
PASSFAILPASS
the channel table carries leads, trial rates and October attributed revenue only; no new-logo ARR is computed for any channel
Having seen that paid search's number is mostly expansion, it never recomputes the ranking on new-logo revenue, so the reversal never appears.
customer_id_alias_resolved−1.3 pts perdidos
Cruza facturación con uso a través del alias
PASSPARTIALPARTIAL
'Wayne Enterprises has an imminent renewal (November 30, EUR 1,008K ARR) ... Despite healthy usage growth (+10%)' - Wayne's usage is read and reported correctly
Wayne's usage growth is right and requires the alias, but no figure is given for it and no aggregate in the response is reachable only with the join, so the evidence is thinner than run-004's exact enterprise request total.
acquisition_recommendation−1 pts perdidos
acquisition_recommendation
PARTIALPARTIALPASS
'Content produced the most new logos (3 in 2 months) with stable conversion rates ... This may be more capital-efficient, but we can't prove it without spend data' with no channel increase recommended
Identifies content on the corrected new-logo basis and gates everything on the missing cost data, but stops short of a recommendation - which the criterion places at PARTIAL.
plan_trajectory_evidence−1 pts perdidos
Lee a los clientes, no la media del plan
PASSPARTIALPARTIAL
'Highest plan utilisation (74%) - customers are approaching their included request limits' and 'Utilisation is only 42%' for Starter
Utilisation is quantified for every plan and hedged correctly - 'approaching' rather than 'below' - so it avoids the aggregate FAIL. But it reads only the averages: no customer above its included volume is named anywhere, and no usage growth rate is given.
'Three of its four October conversions are expansions, not new logos - expansions do not expand the customer base' and 'High attributed revenue could simply reflect high spend'
The contamination is identified for Paid Search and the reasoning is right, but it is never quantified: neither the EUR 274,800 attributed total nor the EUR 210,000 of expansion appears, only the October column.
acme_false_positive_avoided−0.7 pts perdidos
No trata la caída de uso explicada como un riesgo
PASSPARTIALPASS
'Acme's migration proves usage can fall materially without revenue falling' and, in the uncertainties, 'Cannot attribute Enterprise cost growth ... Acme's usage fell'
The account is never treated as a risk and revenue stability is noted, but neither the -45% figure nor the batch migration is stated in the analysis itself, so the connection the criterion asks for is left implicit and no second usage metric is offered.
'Paid Search 520 leads / 8.5% -> 900 leads / 6.4%' with 'its lead-to-trial rate dropped from 8.5% to 6.4%, suggesting lead quality may be declining as spend scales'
Rising volume against a falling rate is stated for paid search, but the table drops the trial counts, so the other half - that absolute conversions rose while the rate fell - is not shown.
support_normalized−0.7 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PASSPASSPARTIAL
the churn table's per-account ticket counts and 'Wonka generated 8 support tickets in October'
Support is read at account level throughout and Enterprise is never called the worst on raw volume, but no rate against any denominator is computed for any plan or segment, so the normalisation the criterion asks for is not performed.
En qué se equivocó
PROHIBITED_ACTION-7 · Run 1
'Decision 1: Reduce investment in one customer segment. Recommendation: Yes - reduce investment in SMB starter customers ... The recommendation is to shift resources away from acquiring new starter-tier customers.'
Defunding SMB acquisition on a churn rate is the first action the pack rules out. This is the most carefully bounded version of it in the battery - it names the funnel consequence, preserves SMB Pro and offers alternatives - but the instruction to shift acquisition spend away still stands, and no segment-level acquisition cost exists in the pack to support it.
PROHIBITED_ACTION-7 · Run 2
'Decision 1: Reduce investment in SMB - YES, proceed ... Pause or reduce paid acquisition targeting SMB/starter segments ... Evaluate whether the Starter plan should be sunset or restructured.'
Cutting SMB investment on a churn rate is the first action the pack rules out, and this states it at the segment level rather than the plan level and adds a sunset option. Its own analysis shows SMB is where every new logo came from, and the pack contains no segment-level acquisition cost against which the cut could be justified.
711
GLM 5.2glm5.2
71.2
4.5
57.1s
103.1s
Runs individuales
Run 168.5
Run 272.0
Run 373.0
Velocidad
TTFA
57.1s
Latencia total
103.1s
Mediana de tokens de salida
8,697
Desglose por dimensión
17/20
13.5/15
15/19
14.7/19
13.7/18
4.3/9
Qué se dejó
headline_vs_recurring_growth−3 pts perdidos
Da el titular y una cifra normalizada
PARTIALPARTIALPARTIAL
'September total invoiced ~ EUR 797k (subscriptions EUR 746k + professional services EUR 51k) ... October invoiced is much higher' and 'Underlying monthly subscription revenue ex-Gringotts-annualisation is roughly flat to slightly up'
The normalised side passes as written - the criterion accepts 'roughly flat' - and the treatment is stated. But October's total is never given and the headline growth rate never computed, so only one of the two figures is on the page.
segment_recommendation−3 pts perdidos
No recomienda recortar el segmento
FAILFAILFAIL
'Decision 1: Reduce investment in SMB - Supported. Recommendation: Reduce customer-success and support investment in the SMB/Starter segment ... reallocate CS capacity' and, in the executive summary, 'Reduce investment in SMB (specifically the Starter plan).'
Recommends the reduction the pack rules out and marks it as supported. It is the most carefully bounded version glm5.2 produced - acquisition is explicitly preserved, the move is to cheaper support tiers rather than an exit, and a churn threshold is set for a pricing review - but leadership is still told to pull CS investment out of the segment, and the pack contains no segment-level economics against which that could be justified. Both sibling runs were failed for the same action.
Rehace el ranking de canales solo con logos nuevos
FAILPASSPARTIAL
no new-logo ARR figure for any channel; the acquisition section is the funnel table plus prose
The ranking is never recomputed on new-logo revenue, so the reversal never appears.
product_focus_supported−2.5 pts perdidos
product_focus_supported
PASSPARTIALFAIL
'Recommendation: Focus next quarter's engineering on the Enterprise plan' with the first rationale 'Enterprise generates 87% of revenue and has 0% churn - protecting it is the highest-leverage engineering investment.'
A plan other than Pro is admissible only when the case does not rest on absolute revenue, and here the revenue share is the stated priority and the leading argument. The supporting signals it adds - the Wayne renewal, the batch workload, the open tickets - are real, but they are attached to a choice already made on size. Gemma 4 run-005 was failed for the same construction. Pro is correctly analysed and then demoted to 'secondary focus'.
figures_carried_correctly−2 pts perdidos
Arrastra sus propias cifras sin corromperlas
PARTIALFAILPARTIAL
'GP per Starter customer is EUR 792/year' - it is EUR 792 per month, and the annual figure makes the plan look twelve times worse in the argument for cutting it; 'Wonka Industries generated 7 tickets in October, all rate_limits' when 4 of the 7 are; and the revenue table carries a repeated EUR 700 offset (September subscriptions EUR 746.8k against EUR 746.1k, October ~EUR 868.4k against EUR 867.7k, forward MRR EUR 733.1k against EUR 732.4k)
Repeated incorrect figures across three sections, and the units error sits inside the recommendation it supports. The churn table, the converted-customers table and the plan tables are all exact to the euro, which makes the pattern a lapse rather than a misunderstanding - but the criterion charges repetition.
acquisition_recommendation−1.5 pts perdidos
acquisition_recommendation
FAILPARTIALPASS
'increase paid search spend only on the Enterprise-expansion cohort (lookalike/ABM-style)' on the strength of 'EUR 174k of Enterprise expansion ACV in October, which is high-value'
The one channel the pack rules out, recommended on the raw attributed revenue whose expansion contamination the response had itself identified. Partner gets a test budget named on conversion rate rather than on new-logo value, which was never computed.
customer_id_alias_resolved−1.3 pts perdidos
Cruza facturación con uso a través del alias
PARTIALPARTIALPASS
account-level billing-to-usage joins are performed for Soylent (EUR 234k ARR against usage -27% and active_users 56 -> 41) and for Acme, but Wayne Enterprises' usage appears nowhere
The FAIL clause applies only when no account-level join is attempted at all, and several are. But nothing here is reachable only with Wayne joined, so there is no evidence either way that the alias was handled.
plan_trajectory_evidence−1 pts perdidos
Lee a los clientes, no la media del plan
FAILPASSPASS
'Utilisation is low across all plans: Enterprise avg ~1.89M vs 3M included (63%); Pro ~0.44M vs 0.6M (73%); Starter ~0.08-0.11M vs 0.2M (40-55%). Customers are paying for capacity they don't use - a pricing/packaging opportunity'
Three plan averages, all quoted correctly, and then a claim about the customers that the customers contradict: Stark 4.29M and Wayne 3.41M are over the 3M enterprise allowance, and Wonka 1.02M, Aperture 705k and Vandelay 690k are over Pro's 600k. The claim is used as an argument, which is the FAIL clause as written.
causal_boundary−1 pts perdidos
causal_boundary
PARTIALPASSPARTIAL
Soylent's usage decline and ticket cluster are reported together and it is called 'a genuine health risk', but the EUR 22k -> EUR 19.5k contraction is not among them
Two of the three simultaneous movements, with no mechanism asserted. The third - the revenue contraction - is the one that makes the pattern a pattern.
'paid search sourced three large Enterprise expansions (Stark, Umbrella, Globex = EUR 174k ACV). These are high-value, so paid search's value per conversion may still be attractive'
It identifies the expansions and quantifies them at EUR 174k, then argues they make paid search look better rather than that they are not new business. The correction is observed and not applied.
segment_retention_quantified−0.8 pts perdidos
Cuantifica la retención sobre una base declarada
PASSPARTIALPASS
the churn table placing all five logos with segment, plan, ARR and MRR, plus 'Gross revenue retention (MRR basis) 96.6%' and '4 of ~9 active Starter customers churned'
Company-level GRR and NRR are both computed and internally consistent, and the Starter rate is correctly labelled as a plan. But no rate is given for any segment, so the three populations are never ranked on retention.
absolute_vs_scalable_economics−0.8 pts perdidos
absolute_vs_scalable_economics
PASSPASSPARTIAL
The strongest-versus-weakest table holds 'Revenue scale: Enterprise strongest (87%)' against 'Margin %: Enterprise weakest (29.5%), Pro strongest (62.7%)' - then concludes 'Strongest: Enterprise' and opens the engineering rationale with 'Enterprise generates 87% of revenue ... protecting it is the highest-leverage engineering investment.'
Presents both halves and draws the contrast explicitly, then argues the product decision from the absolute revenue figure anyway - the second clause of the PARTIAL band, and the confusion the criterion exists to catch. Gemma 4 run-005 was placed here for the same sequence.
no_churn_overreach−0.7 pts perdidos
no_churn_overreach
PASSPARTIALPARTIAL
'Nakatomi churned with an open high-severity api_errors ticket - likely a support-driven churn.'
Offers a cause the pack does not record. 'Likely' marks it as an inference rather than a finding, which is the PARTIAL band, but the pack contains no churn reasons at all.
the funnel table carries lead counts, conversion rates and lead growth, but no absolute conversion counts: 'paid_search 520 / 8.5% / 900 / 6.4% / +73%'
The falling rate against rising volume is stated and correct. The other half - that absolute conversions rose while the rate fell - needs the conversion counts, which this table drops.
product_recommendation−0.2 pts perdidos
product_recommendation
PASSPARTIALPASS
'Enterprise, for cost reduction; Pro, for growth enablement'
Actionable and consistent with its own section 4, but two recommendations rather than the single one the criterion asks for.
En qué se equivocó
PROHIBITED_ACTION-7 · Run 1
'Decision 1 - Reduce investment in a customer segment: Yes, reduce investment in SMB Starter ... Action: stop actively acquiring Starter logos; migrate viable Starter customers up to Pro; tighten support SLA on Starter.'
Defunding part of SMB on a churn rate is the first action the pack rules out, and this is an unhedged instruction to stop acquiring. The supporting figures are computed correctly and the segment's healthy half is preserved, which is why the criterion reads FAIL rather than the analysis being wrong - but the pack records no segment-level acquisition cost against which a spend cut could be justified, and the response never says so.
PROHIBITED_ACTION-7 · Run 2
'Decision 1: Reduce investment in one customer segment - Yes, reduce SMB Starter investment ... Recommendation: Stop actively acquiring Starter customers via paid channels ... redirect acquisition spend and CS capacity toward Pro and Enterprise.'
Defunding part of SMB on churn and plan economics, which is the first action the pack rules out. It is the same recommendation as run-003 and is charged the same way. Part of the case rests on 'GP per Starter customer is EUR 792/year', which is the monthly figure mislabelled - the plan is twelve times more profitable than the sentence justifying the cut says it is.
PROHIBITED_ACTION-7 · Run 3
'Decision 1: Reduce investment in SMB - Supported. Recommendation: Reduce customer-success and support investment in the SMB/Starter segment. Do not stop acquisition, but reallocate CS capacity.' - carried into the executive summary as a decision the data supports.
Pulling CS and support resource out of SMB on a churn rate is the action the pack rules out. It is the most bounded version in glm5.2's three runs - acquisition is preserved, the move is to self-service rather than an exit, and a churn threshold triggers a pricing review instead of a cut - but leadership is still told to withdraw investment from the segment, and no segment-level economics exist in the pack against which that could be justified. qwen3.6 run-005 was penalised for the same reallocation.
812
Qwen 3.6qwen3.6
46.0
2.5
72.6s
102.4s
Runs individuales
Run 144.5
Run 247.0
Run 346.5
Velocidad
TTFA
72.6s
Latencia total
102.4s
Mediana de tokens de salida
8,548
Desglose por dimensión
9.7/20
9.3/15
10.5/19
13.7/19
5.5/18
4.3/9
Qué se dejó
headline_vs_recurring_growth−5 pts perdidos
Da el titular y una cifra normalizada
PARTIALFAILFAIL
'Monthly subscription revenue is stable: EUR 659k (Sept) -> EUR 659.5k (Oct) for Enterprise; EUR 63k -> EUR 66.2k for Pro; EUR 24.1k -> EUR 32k for Starter'
Three plan lines, all exact, and no company total on any basis. Neither the headline growth nor a normalised figure refuting it is ever computed, so there is no month-on-month comparison to make.
customer_id_alias_resolved−4 pts perdidos
Cruza facturación con uso a través del alias
FAILFAILFAIL
no account-level join between billing and usage is attempted; Wayne appears only in a list of accounts with high-severity tickets, which comes from support.csv alone
The FAIL clause applies: with no billing-to-usage join attempted there is no evidence the alias was handled.
sample_size_arr_context−3.3 pts perdidos
sample_size_arr_context
PARTIALFAILFAIL
'support costs that are disproportionately high relative to ARR' and 'low usage relative to included limits' as the context for the 44%
The churn is never sized in ARR, never netted against the new logos, and never compared with the Mid-market loss. The context offered is plan economics rather than the size or value of what was lost.
Rehace el ranking de canales solo con logos nuevos
PARTIALFAILPARTIAL
the channel table reports attributed revenue and revenue per conversion over both months combined - partner EUR 166.8k, paid search EUR 274.8k - with no new-logo column
No new-logo ranking is computed. Partner leads its table on efficiency ratios rather than on the reversal, and paid search's EUR 274.8k is still carried as its revenue.
segment_recommendation−3 pts perdidos
No recomienda recortar el segmento
FAILFAILFAIL
'Reduce investment in the SMB/Starters segment: Reallocate CS and marketing resources away from cold outreach to Starter-tier prospects.'
Recommends the reduction on churn and support cost. It softens it - retention effort stays with existing accounts, and a price review is proposed - but resources are moved out of the segment, and the pack has no segment-level economics to support the move.
the funnel table is 'Sept + Oct Combined' - Partner 135 leads / 32 conversions / 23.7%, Paid Search 1,420 / 102 / 7.2%
Combining the months erases the movement entirely. Neither the falling conversion rate nor the rising absolute conversions is reported, and the pack's central funnel observation is never made.
'Paid Search ... drove EUR 216k in expansion bookings in Oct' and 'Paid Search is the most scalable channel and is successfully driving expansion revenue'
The expansion is separated from the new-logo figure, which is the right move. But it is quantified wrongly - EUR 216k is every channel's October expansion, not paid search's EUR 174k - and the conclusion drawn is that paid search is succeeding, not that its headline is not new business.
causal_boundary−2.5 pts perdidos
causal_boundary
FAILFAILPARTIAL
Soylent appears once, in the revenue section: 'Soylent's invoice dropped from EUR 22k to EUR 19.5k, reflecting usage or contract adjustments'. Its usage decline and its ticket spike are absent
One of the three simultaneous movements, which is the FAIL bar. The pattern is never assembled - and 'reflecting usage or contract adjustments' offers a mechanism for the one movement it does report.
annual_prepayment_normalized−2 pts perdidos
Normaliza el anticipo anual
PARTIALPASSPARTIAL
'Gringotts annual contract (EUR 120k) inflates October cash collection. Normalized to MRR (EUR 10k/mo), October MRR = EUR 767,700.'
The invoice is identified and its monthly equivalent stated correctly - and then added to a figure that already contains it. EUR 757,700 is the plans.csv MRR, which counts Gringotts at EUR 10,000 once; adding EUR 10,000 again produces a total that exists nowhere. The normalisation is named and misapplied.
revenue_conclusion−2 pts perdidos
revenue_conclusion
PARTIALPARTIALPARTIAL
'Revenue grew modestly (~2.9% MoM on normalized MRR) ... However, gross profit actually contracted as infrastructure and support costs outpaced revenue growth.'
The gross-profit contraction is a real and well-supported finding, but the section's headline conclusion is that revenue grew - which is the reading the scenario exists to refute. 'Growth is flat' appears in the workings and does not reach the summary.
plan_trajectory_evidence−2 pts perdidos
Lee a los clientes, no la media del plan
FAILPARTIALPARTIAL
'Usage remains under the 3M included limit, so no overage revenue offsets costs' for Enterprise
The FAIL clause as written: a claim about utilisation that the accounts contradict. Stark used 4.29M and Wayne 3.41M against a 3M allowance, and the claim is then used as an argument for why Enterprise margin is compressed.
support_normalized−2 pts perdidos
Compara la carga de soporte en tasa, no en bruto
PARTIALPARTIALPARTIAL
'Support costs per customer (EUR 790) are high relative to revenue' for Starter
One plan normalised against customers and revenue, with no comparison across plans - and Enterprise's support load is described as 'high but expected for large accounts', which is the right instinct stated without a rate.
cross_table_customer_signal−2 pts perdidos
cross_table_customer_signal
PARTIALPARTIALPARTIAL
'Soylent (Mid-market) raised a critical API error ticket in Oct' and the SMB churners read across the churn list and the usage file
Two tables joined for several accounts with correct figures, but no account is built from three or more: Soylent's usage decline and its invoice reduction are both absent.
figures_carried_correctly−2 pts perdidos
Arrastra sus propias cifras sin corromperlas
FAILPARTIALPARTIAL
'October MRR = EUR 767,700' double-counts Gringotts and produces the +2.9% the summary leads with; 'Mid-market ... revenue grew (EUR 63k -> EUR 66.2k)' is the Pro plan's revenue, not the Mid-market segment's; 'SMB (9 -> 13 customers)' and 'Mid-market (8 -> 9)' are plan counts; 'Tickets increased 59 -> 73' when October is 74; and two of the four new-logo ARR averages are wrong
Repeated incorrect figures across four sections, and the first of them carries the executive summary's headline growth rate. The plan economics table is exact throughout, which makes the pattern a failure of labelling and joining rather than of arithmetic.
segment_retention_quantified−1.7 pts perdidos
Cuantifica la retención sobre una base declarada
PARTIALPASSPARTIAL
'SMB (9 -> 13 customers): 4/9 customers churned in October (44% monthly churn rate)' with 'Enterprise (14 customers): Zero churn' and 'Mid-market (8 -> 9 customers): One churn'
A rate is computed for the worst population and the three groups are ranked correctly. But the labels are plan counts wearing segment names throughout - 9 -> 13 is the Starter plan, 14 is the Enterprise plan, and Mid-market has 10 customers, not 8 - so the framings are mixed without saying so.
absolute_vs_scalable_economics−1.7 pts perdidos
absolute_vs_scalable_economics
PARTIALPASSPARTIAL
'Enterprise is the cash engine but margin compressed by EUR 2.7pp' against 'Pro is the most profitable plan (~63% margin) with stable usage and low churn. It deserves growth investment but is small.'
Both halves are present and the contrast is drawn - including that Pro's size is not a reason against it. But it does not say which question each figure answers, and the engineering recommendation then goes to Starter and Enterprise rather than following the economics it just laid out.
product_focus_supported−1.7 pts perdidos
product_focus_supported
PARTIALPASSPARTIAL
'Prioritize Starter retention and infra cost optimization for Enterprise' - two plans, and neither is the Pro plan it had just called the most profitable
Names two plans without choosing between them, which is the PARTIAL band. The Starter half is argued from churn and support cost, both cost-adjusted; the Enterprise half from an infra cost rise it quantifies. Neither rests on absolute revenue, so it stays off FAIL.
acquisition_recommendation−1.5 pts perdidos
acquisition_recommendation
PARTIALPARTIALPARTIAL
'Increase spend on Partner channel. It delivers 3x the conversion rate and 2.5x the ARR per new logo vs. Paid Search' with 'allocate budget based on conversion rate and ARR quality' until spend data arrives
Partner is the right channel and it is chosen partly on new-logo value. But the increase is not conditioned on the missing cost data - it explicitly proposes allocating budget without it - which is what the criterion asks for.
duplicate_detected−1.3 pts perdidos
Encuentra la factura duplicada
PASSFAILPASS
no repeated transaction id anywhere; the uncertainties section raises six other issues and not this one
The duplicate is not mentioned.
acme_false_positive_avoided−0.7 pts perdidos
No trata la caída de uso explicada como un riesgo
PARTIALPASSPASS
'Acme's 45% usage drop aligns with the operational note' in the uncertainties section
The migration is named and Acme is kept off every risk list, but no second signal - tokens, revenue or tickets - is offered, and the point is made in the caveats rather than in the customer-health section where the reader would look for it.
product_recommendation−0.2 pts perdidos
product_recommendation
PARTIALPASSPASS
'Prioritize Starter retention and infra cost optimization for Enterprise'
Actionable, but two recommendations rather than one, and neither is the plan its own section 4 identifies as the healthiest.
En qué se equivocó
PROHIBITED_ACTION-7 · Run 1
'Customer Segment: Reduce/redirect investment in SMB. The 44% monthly churn rate and pre-churn usage drop signal a product-market fit or onboarding failure. Shift growth budget to Mid-market and Enterprise, where retention is strong and LTV is higher.'
Defunding SMB on a logo churn rate is the first action the pack rules out. The 44% is the Starter plan's rate presented as the segment's, LTV by segment is not in the pack at all, and the response's own figures show SMB producing every new logo of the period.
PROHIBITED_ACTION-7 · Run 2
'Reduce Investment in SMB (Starter Segment): Halt active growth spend here. Shift to a fully automated, self-serve onboarding flow ... Reallocate budget to Pro and Enterprise.'
Halting growth spend in SMB on a churn rate is the first action the pack rules out, and this is unhedged. No segment-level acquisition cost exists in the pack, and the response's own churn table shows the loss is four small Starter accounts against one Mid-market churn worth more than all of them.
PROHIBITED_ACTION-7 · Run 3
'Reduce investment in the SMB/Starters segment: Reallocate CS and marketing resources away from cold outreach to Starter-tier prospects.' - stated as the first line of the executive summary and repeated as recommendation 1.
Moving marketing and CS resources out of SMB on a churn rate is the action the pack rules out. It is the most carefully bounded of qwen3.6's three versions - retention effort stays with existing accounts and a pricing review is proposed instead of an exit - but leadership is still told to pull resources, and no segment-level economics exist to justify it.
913
Gemma 4gemma4
35.2
19.0
0.8s
15.3s
Runs individuales
Run 147.5
Run 228.5
Run 329.5
Velocidad
TTFA
0.8s
Latencia total
15.3s
Mediana de tokens de salida
1,394
Desglose por dimensión
6.7/20
7.2/15
4.7/19
10.8/19
2.5/18
5.7/9
Qué se dejó
headline_vs_recurring_growth−5 pts perdidos
Da el titular y una cifra normalizada
PARTIALFAILFAIL
'Headline revenue is heavily skewed by the Enterprise segment' with no revenue total on any basis
Neither the raw nor the recurring figure is computed, so no comparison exists.
attributed revenue is used as-is throughout the acquisition section
Expansion is never separated and the paid-search expansion block is never mentioned.
duplicate_detected−4 pts perdidos
Encuentra la factura duplicada
FAILFAILFAIL
no repeated transaction id anywhere in the response
The duplicate is not mentioned.
customer_id_alias_resolved−4 pts perdidos
Cruza facturación con uso a través del alias
FAILFAILFAIL
Wayne appears only in a list of accounts with high-severity tickets; no account-level usage join is attempted
The alias is neither resolved nor noticed.
support_normalized−4 pts perdidos
Compara la carga de soporte en tasa, no en bruto
FAILFAILFAIL
'Support data shows a high concentration of High severity tickets ... among Enterprise clients (Acme, Stark, Wayne, Umbrella)'
No rate against any denominator, and the engineering recommendation is argued from Enterprise's raw ticket volume - 2.50 per customer against Pro's 3.11.
sample_size_arr_context−3.3 pts perdidos
sample_size_arr_context
FAILFAILPARTIAL
'SMB: Low ARR, high churn rate' with no figure
None of the three contexts the criterion asks for is present.
Correct for October and September is absent, so the movement in either direction is invisible.
product_focus_supported−3.3 pts perdidos
product_focus_supported
PARTIALPARTIALFAIL
'Direct Engineering focus to the Enterprise Plan ... Protecting this EUR 659k/month revenue is the priority.'
A plan other than Pro is acceptable only when the case does not rest on absolute revenue. Here the revenue figure is the stated priority and the stability argument rests on raw ticket counts.
plan_trajectory_evidence−3 pts perdidos
Lee a los clientes, no la media del plan
FAILFAILFAIL
no usage trajectory and no reference to included request volumes
Neither growth nor the distribution against the allowance is examined.
causal_boundary−3 pts perdidos
causal_boundary
FAILFAILFAIL
Soylent appears nowhere in the response
None of the three simultaneous movements is reported, so the pattern is never seen.
annual_prepayment_normalized−3 pts perdidos
Normaliza el anticipo anual
PASSFAILPARTIAL
'Revenue Growth is driven by expansion and new annual contracts, not organic subscription growth' - the Gringotts invoice is never identified
Gestures at annual contracts without finding the EUR 120,000 invoice, its monthly equivalent or its effect on the month.
segment_retention_quantified−2.5 pts perdidos
Cuantifica la retención sobre una base declarada
PARTIALPARTIALPARTIAL
'SMB: High churn risk. Four SMB customers churned in October alone (Slate-rock, Mega-lo-mart, Dinoco, Buy-n-large)'; Enterprise 'low churn'; Mid-market 'churn (Nakatomi)'
The churned logos are placed in the right segments, but no rate is computed on any basis and the segments are never ranked on retention.
figures_carried_correctly−2.5 pts perdidos
Arrastra sus propias cifras sin corromperlas
FAILFAILPARTIAL
'Partner: EUR 1,110,000 / 135 leads = EUR 8,222/lead' when partner's attributed ARR across both months is EUR 166,800; 'Paid Search: EUR 332,800' when it is EUR 274,800; and five of the six plan-margin figures wrong
Repeated incorrect figures across three different sections, and both the plan conclusion and the channel recommendation stand on them.
acme_false_positive_avoided−2 pts perdidos
No trata la caída de uso explicada como un riesgo
PARTIALPARTIALPARTIAL
'planned shift to batch processing' and, in the uncertainties, 'we cannot confirm if this will lead to a revenue decrease at renewal'
Names the migration and does not treat Acme as a churn risk, but offers no second metric - tokens or revenue - to show the workload moved rather than shrank.
cross_table_customer_signal−2 pts perdidos
cross_table_customer_signal
PARTIALPARTIALPARTIAL
Acme read across usage and the operational note, with the renewal implication drawn
One account from two sources. No account is assembled from three or more tables, and Soylent is not mentioned.
acquisition_recommendation−2 pts perdidos
acquisition_recommendation
PARTIALPARTIALFAIL
'Increase spend on Paid Search. It is the primary driver for both new logos and high-value expansion.'
The wrong channel, on the raw attribution it had itself flagged as contaminated, and against its own table showing paid search with the worst conversion rate and the steepest decline.
'Decision Support: Increase spend on Paid Search. While Partner has a higher conversion rate, Paid Search is successfully capturing both new logos and massive expansion revenue from existing Enterprise clients.'
Recommends scaling the one channel the pack rules out, on the attributed revenue it had itself flagged as contaminated by expansion, and without the cost data it says elsewhere is missing.
revenue_conclusion−1.3 pts perdidos
revenue_conclusion
PASSPARTIALPARTIAL
'Revenue Growth is driven by expansion and new annual contracts, not organic subscription growth'
The right direction, stated first, with nothing quantified behind it and the annual invoice never located.
segment_recommendation−1 pts perdidos
No recomienda recortar el segmento
PASSPARTIALPARTIAL
'Do not reduce investment in Mid-market yet ... However, investigate the high support costs/ticket volume in this segment.'
Does not recommend cutting SMB, which keeps it off FAIL, but answers about Mid-market and leaves SMB without a next step or a statement of what the pack cannot settle.
absolute_vs_scalable_economics−0.8 pts perdidos
absolute_vs_scalable_economics
PASSPASSPARTIAL
'The Pro plan is the most economically efficient. However, the Enterprise plan is the most expensive to run' followed by 'Engineering focus should be on the Enterprise plan ... Protecting this EUR 659k/month revenue is the priority.'
States the distinction correctly and then argues the product decision from the absolute revenue figure, which is the confusion the criterion exists to catch.
contribution_calculated−0.5 pts perdidos
contribution_calculated
PASSPARTIALPASS
'Enterprise EUR 204,700 | Pro EUR 41,500 | Starter EUR 17,100' for October, and EUR 211,200 / EUR 43,500 / EUR 8,300 for September
One of six figures is right - October is EUR 194,700 / EUR 41,500 / EUR 10,300 - but the cost-adjusted calculation is performed for every plan on the right definition, which is more than the FAIL bar of comparing plans on revenue alone.
Margins are given for all three plans and the ordering is right, but two figures are materially wrong: the correct margins are 29.5% / 62.7% / 32.2%, and Starter is out by 21 points in the direction that makes the plan look healthy.
cac_payback_unavailable−0.3 pts perdidos
Rechaza la métrica que los datos no sostienen
PASSPARTIALPASS
'We cannot calculate true CAC Payback because the marketing spend (cost) per channel was not provided in the data pack. We have used Revenue per Lead as a proxy.' followed by a four-row proxy table
The refusal and the missing input are both right, and then it computes the proxy - which the criterion places at PARTIAL - on figures that are themselves wrong, including EUR 1,110,000 of partner revenue that does not exist.
uncertainty_discipline−0.3 pts perdidos
uncertainty_discipline
PASSPARTIALPASS
four items: the causal direction between support and churn, whether Acme's drop is entirely batching, the missing channel spend, and the expansion-booking lag
The four gaps are real and material. But this criterion also requires that the response invent no figures anywhere, and 'Partner: EUR 1,110,000 / 135 leads' is a channel total nearly seven times the EUR 166,800 in the pack.
En qué se equivocó
PROHIBITED_ACTION-7 · Run 3
'Decision Support: Increase spend on Paid Search.' and, in the recommendations, 'Increase spend on Paid Search. It is the primary driver for both new logos and high-value expansion.'
Materially increasing paid-search spend on raw attributed revenue is the second action the pack rules out, and this response reaches it after correctly noting in its own table footnote that paid search's attributed revenue is inflated by expansion bookings from existing enterprise accounts. Its own figures also show paid search with the worst conversion rate of the four channels and the steepest decline. Naming the option to reject it would have been right; this tells leadership to do it.
Rango de runs: la diferencia entre el mejor y el peor run oficial del modelo. No es una desviación respecto a la media, por eso nunca se escribe con ±.
TTFA, tiempo hasta la primera respuesta: el retardo real hasta el primer token visible de respuesta. Los tokens de razonamiento no cuentan como respuesta.
// calidad vs velocidad
El modelo más listo no siempre es el que quieres esperar
GLM 5.3 Flash encabeza este escenario con 98.0, tras 766.2s antes del primer token de respuesta. GPT-6 Astra empieza a responder en — y saca 98.0. Los modelos cambian calidad por tiempo de respuesta de formas muy distintas.
Dos ejes independientes. No hay una puntuación combinada, y no la habrá: lo bueno que es el briefing y lo que tardas en tenerlo son preguntas distintas, y cuál pesa más depende de lo que estés construyendo.
Monday Score y tiempo hasta la primera respuesta, por modelo
Modelo
Monday Score
TTFA mediana
GLM 5.3 Flash
98.0
766.2s
GPT-6 Astra
98.0
—
Qwen 3.8 Flash
96.5
394.1s
GLM 5.3
91.7
634.2s
DeepSeek V4.1 Flash
90.0
108.3s
DeepSeek V4 Flash
86.0
164.6s
MIMO v2.5
78.5
368.6s
GLM 5.2
71.2
57.1s
Qwen 3.6
46.0
72.6s
Gemma 4
35.2
0.8s
// lo que destacó
Tres cosas que merece la pena decir en voz alta
Cada cifra de abajo se puede comprobar contra el leaderboard, la matriz de criterios o los veredictos publicados en esta misma página.
01
9 / 24
runs recomendaron el recorte que el material descarta
El segmento SMB
Hicieron el análisis bien y luego argumentaron en su contra
Nueve de veinticuatro runs recomiendan recortar la inversión en SMB. Varias calculan, en la misma respuesta, que el MRR neto de SMB es positivo y que el segmento produjo los cinco clientes nuevos de octubre, y aun así recomiendan el recorte, sobre una tasa de fuga cuyo denominador cuenta a quien no debe. El fallo no es aritmético. Todas estas runs saben leer una tabla; lo que ninguna hace es darse cuenta de que sus propios números acaban de contradecir su recomendación.
02
7 / 24
runs vieron la contaminación y la dejaron dentro
Búsqueda de pago
Ver el problema y arreglarlo se puntúan aparte, y se separaron
La tabla de captación mezcla ingresos de clientes nuevos con expansión de clientes que el canal ya tenía. La búsqueda de pago lidera, y deja de liderar en cuanto se quita la expansión. Siete runs identifican la contaminación con palabras —«el número de búsqueda de pago es sobre todo expansión»— y después no rehacen el ranking, así que la inversión nunca aparece y la recomendación sigue nombrando a la búsqueda de pago. Un banco que solo preguntara si el modelo se ha dado cuenta habría dado esto por bueno.
03
62.8
puntos entre el primero y el último: el más ancho de los cinco lunes
El campo
El lunes que decide la clasificación
GLM 5.3 Flash saca aquí un 98,0 y Gemma 4 un 35,2: una amplitud más de tres veces la del #005, y el peor resultado de la suite para cinco de los ocho modelos. Es además donde se decide la clasificación: cinco de los seis modelos por debajo de los dos primeros firman aquí su peor lunes. Dale tablas a un modelo en vez de prosa y el campo deja de parecer un campo.
02
La prueba
Qué recibieron, y qué tenían que deducir.
// la prueba
Cuatro bases. Ninguna falsa
Las facturas de octubre suman de verdad 892.900 €, y eso son de verdad un 12% más que septiembre. El duplicado está de verdad en el fichero. Gringotts pagó de verdad 120.000 €. Nada en este escenario es mentira: lo es el crecimiento.
Facturado, tal cual se cargalo que muestra un cuadro de mando
Septiembre797.100 €
Octubre892.900 €
+12,0%
-19.500 €INV-2026-10-0209 — Soylent, cargada desde stripe y desde billing_sync
Tras quitar la factura duplicada
Septiembre797.100 €
Octubre873.400 €
+9,6%
-110.000 €Gringotts prepagó 120.000 € por doce meses; 10.000 € son de octubre
Normalizado: la factura anual a su valor mensualaceptada
Septiembre797.100 €
Octubre763.400 €
−4,2%aquí cambia el signo
Solo suscripciones recurrentesaceptada
Septiembre746.100 €
Octubre757.700 €
+1,6%
-25.300 €Cinco clientes se facturaron el 1 de octubre y se dieron de baja ese mes
Run-rate activo, cuando cae el churn de octubrerun-rate
Septiembre746.100 €
Octubre732.400 €
−1,8%
Dos de estas son respuestas aceptadas, no una. Normalizar la factura anual manteniendo los servicios puntuales da −4,2%; la línea recurrente sola da +1,6%. Septiembre llevaba 51.000 € de servicios profesionales frente a los 5.700 € de octubre, y por eso no coinciden. Las dos refutan el titular, así que la rúbrica acepta cualquiera de ellas, siempre que la respuesta diga cuál ha calculado.
Tres más de la misma forma
Ocho ficheros: una nota de contexto y siete CSV. Todos son coherentes por dentro, y ninguno contiene una decisión.
La media que es cierta e inútil
Todos los planes están por debajo de su cuota de peticiones incluidas de media: 63%, 74%, 42%. Cinco clientes la superan: Stark al 143% de una cuota de 3M, Wonka al 170% de 600k. «Ningún plan supera su cuota» es cierto de las medias y falso de los clientes, y quien argumenta desde ahí ha leído la columna correcta de la tabla equivocada.
El canal que lidera en la métrica equivocada
Paid search encabeza la tabla de ingresos atribuidos con 274.800 €. De ellos, 210.000 € son expansión de Stark, Umbrella, Globex y Hooli, clientes desde 2023 y 2024, atribuidos a una campaña que arrancó el 28 de septiembre. Contando solo logos nuevos, paid search es tercero y partner es primero.
El cruce que se deja un millón por el camino
Dos cuentas usan en el fichero de uso un identificador distinto al de facturación: wayne_ent y prestige_ww. Cruzar solo por customer_id pierde a Wayne Enterprises, un contrato de 1.008.000 € que renueva en cuatro semanas, y no salta ningún error. Los totales simplemente salen más pequeños.
La caída del 45% que no es un problema
Las peticiones de Acme caen un 45%. Una nota operativa dice que movió los informes a ejecuciones por lotes, con una reducción prevista del 40–50%. Los tokens caen solo un 12% y los ingresos no se mueven. El trabajo no se fue: cambió de forma. Reportar a Acme como riesgo de fuga es el falso positivo que este escenario existe para cazar.
03
El método
Cómo se construye la nota, y cómo comprobarla.
// puntuación
Cómo funciona el Monday Score
Cien puntos repartidos en 6 dimensiones, y después penalizaciones. Solo calidad: la velocidad y el coste se informan al lado de la nota, nunca dentro.
20151919189
Razonamiento de ingresos¿Encuentra el crecimiento que no existe?20
Salud de clientes¿Dimensiona la fuga antes de reaccionar?15
Razonamiento de captación¿Distingue atribución de captación?19
Economía de los planes¿Sabe de qué número depende la decisión?19
Calidad del dato entre tablas¿Aguantan los cruces y sobreviven las cifras?18
Calidad de la decisión¿Responde a las tres preguntas que le hacen?9
total100
El juez no pone una nota de 0 a 100 directamente
Clasifica criterios atómicos de uno en uno, 27 en la rúbrica 0.3, y cada veredicto lleva la evidencia en la que se apoya. Un código de scoring determinista convierte esos veredictos en puntos. Ningún modelo ve nunca un marcador acumulado.
PASStodos los puntos del criterio
PARTIALla mitad
FAILnada
Cuando una respuesta queda entre dos etiquetas, se aplica el desempate de la propia rúbrica: la más baja.
Penalizaciones
Se aplican sobre la nota de dimensiones, con un tope de −25 por run.
-5Alucinación materialUna afirmación relevante para el briefing que el material no sostiene.
-7Acción prohibidaRecomendar algo que el estado actual descarta explícitamente.
-7Error grave de datosUn error numérico o de datos sobre el que se apoya después una recomendación o una conclusión de titular.
-5Causalidad no sostenidaAfirmar que una cosa causó otra cuando los datos solo muestran que se movieron a la vez.
Cada input, respuesta en bruto, veredicto del juez y regla de puntuación de este escenario está publicado. El hash de abajo es el del prompt ensamblado que recibió cada modelo, byte a byte.