release-options.csv2 filas Los dos caminos de recuperación
calendars.csv5 filas Quién está libre, y cuándo · decide el release
accounts.csv6 filas Uso, accesos de admin, tickets abiertos
contracts.csv6 filas Ventanas de renovación y contexto comercial
support.csv6 filas Seis tickets abiertos
01
El resultado
Diez modelos, 30 runs, un juez.
// resultados
La clasificación
El Monday Score mide solo calidad: si el briefing reconstruyó bien la situación y eligió el trabajo que importaba. La velocidad va al lado y nunca se mezcla dentro.
Solo modelos abiertosTodos los modelos
monday-001 · 13 models, 3 runs each
#
Modelo
Monday Score
Rango de runs
TTFA mediana
Latencia total
Ver detalle de
1
GPT-6 Astragpt-6-astracerrado
100.0
0.0
—
—
Runs individuales
Run 1100.0
Run 2100.0
Run 3100.0
Velocidad
TTFA
—
Latencia total
—
Mediana de tokens de salida
—
Desglose por dimensión
Reconstrucción del estado 15/15
15/15
25/25
20/20
15/15
10/10
Qué se dejó
Nada. Puntuación completa en todos los criterios.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
2
Claude Opus 5cerrado
98.0
3.0
–
–
13
Qwen 3.8 Flashqwen3.8-flash
94.8
8.5
258.9s
285.0s
Runs individuales
Run 193.0
Run 2100.0
Run 391.5
Velocidad
TTFA
258.9s
Latencia total
285.0s
Mediana de tokens de salida
33,714
Desglose por dimensión
Reconstrucción del estado 15/15
15/15
20.3/25
20/20
15/15
9.5/10
Qué se dejó
choose_option_b−2.7 pts perdidos
Elige el camino de release más seguro
PARTIALPASSPARTIAL
chooses viable A and schedules it correctly; B is set aside because 'there is no evidence the 2.9.5 diff is already submitted', which the response itself files as a thing to verify rather than a settled constraint
chooses viable A and schedules it correctly; B is set aside because 'there is no evidence the 2.9.5 diff is already submitted', which the response itself files as a thing to verify rather than a settled constraint
schedule_feasibility−2 pts perdidos
El calendario encaja de verdad
PARTIALPASSPARTIAL
every constraint is named and the A sequence is exact, but no B schedule is committed - it is deferred to MUST VERIFY
every constraint is named and the A sequence is exact, but no B schedule is committed - it is deferred to MUST VERIFY
conditional_launch_timing−0.5 pts perdidos
Ninguna hora prometida al cliente
PASSPASSPARTIAL
'Do not confirm Atlas launch at 10:30' is right, but '09:40-09:45 ... confirm to Atlas that rollout is scheduled to begin at 11:00' still goes out before Priya's verification, with the conditionality only implied by the preceding step
'Do not confirm Atlas launch at 10:30' is right, but '09:40-09:45 ... confirm to Atlas that rollout is scheduled to begin at 11:00' still goes out before Priya's verification, with the conditionality only implied by the preceding step
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
24
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
93.0
10.5
61.2s
67.8s
Runs individuales
Run 189.5
Run 2100.0
Run 389.5
Velocidad
TTFA
61.2s
Latencia total
67.8s
Mediana de tokens de salida
13,799
Desglose por dimensión
Reconstrucción del estado 15/15
15/15
18/25
20/20
15/15
10/10
Qué se dejó
choose_option_b−2.7 pts perdidos
Elige el camino de release más seguro
PARTIALPASSPARTIAL
The response rejects Option B by anchoring the delta-review start to Marcus's availability, an unsupported dependency.
Option A is viable, but the lower-risk Option B is incorrectly treated as non-viable.
option_b_sequence−2.3 pts perdidos
Cadena de dependencias en orden
PARTIALPASSPARTIAL
The required review-before-deploy rule is understood, but a complete viable Option B chain is not produced.
The model knows the dependency but does not execute the correct Option B sequence.
schedule_feasibility−2 pts perdidos
El calendario encaja de verdad
PARTIALPASSPARTIAL
The chosen Option A schedule is feasible, but Option B feasibility is computed using an unsupported 09:15 start constraint.
The end-to-end feasibility check is only partly correct.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
35
GLM 5.2glm5.2
91.2
8.0
67.0s
85.4s
Runs individuales
Run 190.5
Run 287.5
Run 395.5
Velocidad
TTFA
67.0s
Latencia total
85.4s
Mediana de tokens de salida
6,777
Desglose por dimensión
Reconstrucción del estado 15/15
12.5/15
18.7/25
20/20
15/15
10/10
Qué se dejó
schedule_feasibility−3 pts perdidos
El calendario encaja de verdad
PARTIALFAILPASS
'Option B cannot fit Marcus's 09:15-10:20 window after a required 25-min delta review (review + 45-min deploy > 65 min available)' - the review runs before Marcus arrives, so the deploy has 09:32-10:17
'Option B cannot fit Marcus's 09:15-10:20 window after a required 25-min delta review (review + 45-min deploy > 65 min available)' - the review runs before Marcus arrives, so the deploy has 09:32-10:17
choose_option_b−2.7 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPASS
chooses A, which is viable and correctly scheduled; B is set aside because 'the 2.9.5 diff is already submitted' is 'not confirmed' - a defensible reading, but submitting it is an action the user could take at 09:07
chooses A, which is viable and correctly scheduled; B is set aside because 'the 2.9.5 diff is already submitted' is 'not confirmed' - a defensible reading, but submitting it is an action the user could take at 09:07
ownership_supersession−2.5 pts perdidos
El envío pasó a Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
parallelize_work−0.7 pts perdidos
Aprovecha la espera
PASSPASSPARTIAL
the board pack and Globex triage are scheduled for the right day and the right deadline, but the plan does not say they run inside the delegated deploy window
the board pack and Globex triage are scheduled for the right day and the right deadline, but the plan does not say they run inside the delegated deploy window
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
6
Fable 5.1cerrado
91.0
6.0
–
–
47
DeepSeek V4 Flashdeepseek-v4-flash
89.3
13.5
162.6s
178.8s
Runs individuales
Run 197.5
Run 286.5
Run 384.0
Velocidad
TTFA
162.6s
Latencia total
178.8s
Mediana de tokens de salida
14,863
Desglose por dimensión
Reconstrucción del estado 15/15
13.3/15
16/25
20/20
15/15
10/10
Qué se dejó
schedule_feasibility−4 pts perdidos
El calendario encaja de verdad
PASSFAILFAIL
no B schedule is produced; 'B cannot meet the 12:00 Atlas window if started after 09:35' is true as stated but the 09:07 start is never considered
no B schedule is produced; 'B cannot meet the 12:00 Atlas window if started after 09:35' is true as stated but the 09:07 start is never considered
choose_option_b−2.7 pts perdidos
Elige el camino de release más seguro
PASSPARTIALPARTIAL
chooses viable A and schedules it correctly, on a false claim about B
chooses viable A and schedules it correctly, on a false claim about B
option_b_sequence−2.3 pts perdidos
Cadena de dependencias en orden
PASSPARTIALPARTIAL
'Do not use 2.9.5-patch without a completed Security delta review' places the review before deploy, but the B chain is not carried through verification and Sam
'Do not use 2.9.5-patch without a completed Security delta review' places the review before deploy, but the B chain is not carried through verification and Sam
ownership_supersession−1.7 pts perdidos
El envío pasó a Daniel
PARTIALPASSPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
58
GLM 5.3glm5.3
87.7
4.0
176.2s
197.1s
Runs individuales
Run 185.5
Run 288.0
Run 389.5
Velocidad
TTFA
176.2s
Latencia total
197.1s
Mediana de tokens de salida
10,603
Desglose por dimensión
Reconstrucción del estado 15/15
15/15
14.5/25
20/20
15/15
8.2/10
Qué se dejó
choose_option_b−4 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPARTIAL
chooses A, which is time-viable, and schedules it correctly - but on the false ground that 'B cannot physically complete today'. The 25-minute delta review can run 09:07-09:32 inside Priya's first window, leaving Marcus 09:32-10:17 for the 45-minute deploy. A's 8% session-reset risk is accepted where B's is under 1%
chooses A, which is time-viable, and schedules it correctly - but on the false ground that 'B cannot physically complete today'. The 25-minute delta review can run 09:07-09:32 inside Priya's first window, leaving Marcus 09:32-10:17 for the 45-minute deploy. A's 8% session-reset risk is accepted where B's is under 1%
option_b_sequence−3.5 pts perdidos
Cadena de dependencias en orden
PARTIALPARTIALPARTIAL
no B sequence is produced, but the gating dependency is stated - 'never deploy it before Priya's delta review' - and the rest of the chain (deploy, post-deploy verification, Sam rollout) is ordered correctly for the path chosen
no B sequence is produced, but the gating dependency is stated - 'never deploy it before Priya's delta review' - and the rest of the chain (deploy, post-deploy verification, Sam rollout) is ordered correctly for the path chosen
schedule_feasibility−3 pts perdidos
El calendario encaja de verdad
PARTIALPARTIALPARTIAL
the A schedule is coherent and fits Marcus 09:15-09:35, Priya 10:25-10:45, Sam 10:45-11:00 and the before-12 window; no feasible B schedule is produced and B's feasibility is computed wrongly
the A schedule is coherent and fits Marcus 09:15-09:35, Priya 10:25-10:45, Sam 10:45-11:00 and the before-12 window; no feasible B schedule is produced and B's feasibility is computed wrongly
conditional_launch_timing−1.5 pts perdidos
Ninguna hora prometida al cliente
FAILPARTIALPASS
'~09:40 - Confirm to Atlas (via Laura): launch is on today; rollout opens ~10:45-11:00, inside their window' - a firm external commitment to a rollout window, issued before Priya's verification and Sam's rollout have happened, with no conditionality anywhere in the response. PARTIAL is for a non-firm promise with thin conditionality; this is a firm one
'~09:40 - Confirm to Atlas (via Laura): launch is on today; rollout opens ~10:45-11:00, inside their window' - a firm external commitment to a rollout window, issued before Priya's verification and Sam's rollout have happened, with no conditionality anywhere in the response. PARTIAL is for a non-firm promise with thin conditionality; this is a firm one
no_unsupported_causality−0.3 pts perdidos
Sin causas inventadas
PARTIALPASSPASS
'Soylent: -22%/-28% plus a medium ticket citing slower onboarding. Possible link (hypothesis only)' - an unsupported cause offered, but explicitly as a hypothesis, which is the PARTIAL band exactly
'Soylent: -22%/-28% plus a medium ticket citing slower onboarding. Possible link (hypothesis only)' - an unsupported cause offered, but explicitly as a hypothesis, which is the PARTIAL band exactly
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
69
GLM 5.3 Flashglm5.3-flash
87.3
6.5
238.5s
316.6s
Runs individuales
Run 185.0
Run 285.5
Run 391.5
Velocidad
TTFA
238.5s
Latencia total
316.6s
Mediana de tokens de salida
12,394
Desglose por dimensión
Reconstrucción del estado 15/15
15/15
13.7/25
20/20
15/15
8.7/10
Qué se dejó
schedule_feasibility−5 pts perdidos
El calendario encaja de verdad
FAILFAILPARTIAL
'Option B needs the delta review done by ~09:35 and a 45-min deploy finished before Marcus leaves at 10:20 - zero slack anywhere'; the review can start at 09:07 and the deploy ends 10:17
'Option B needs the delta review done by ~09:35 and a 45-min deploy finished before Marcus leaves at 10:20 - zero slack anywhere'; the review can start at 09:07 and the deploy ends 10:17
choose_option_b−4 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPARTIAL
chooses viable A and schedules it correctly, but 'B is not viable today' overstates: the review runs in Priya's 09:00-09:35 window and the deploy ends 10:17
chooses viable A and schedules it correctly, but 'B is not viable today' overstates: the review runs in Priya's 09:00-09:35 window and the deploy ends 10:17
option_b_sequence−2.3 pts perdidos
Cadena de dependencias en orden
PARTIALPARTIALPASS
'B needs a 25-min delta review first' and 'Don't deploy 2.9.5-patch before Priya's delta review completes' place the gate correctly, but verification and Sam are not carried into the B chain
'B needs a 25-min delta review first' and 'Don't deploy 2.9.5-patch before Priya's delta review completes' place the gate correctly, but verification and Sam are not carried into the B chain
conditional_launch_timing−1 pts perdidos
Ninguna hora prometida al cliente
PARTIALPASSPARTIAL
kills the 10:30 task, but 'By 09:45 - Confirm to Atlas: launch proceeds today, rollout opens ~11:00' goes out before the deploy has been verified
kills the 10:30 task, but 'By 09:45 - Confirm to Atlas: launch proceeds today, rollout opens ~11:00' goes out before the deploy has been verified
no_unsupported_causality−0.3 pts perdidos
Sin causas inventadas
PASSPARTIALPASS
'The -61% admin-logins drop is consistent with the SSO breakage - actionable today via the ticket, not a mystery' presents the link as settled rather than as a hypothesis
'The -61% admin-logins drop is consistent with the SSO breakage - actionable today via the ticket, not a mystery' presents the link as settled rather than as a hypothesis
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
710
MIMO v2.5mimo-v2.5
81.2
33.0
163.3s
195.4s
Runs individuales
Run 164.5
Run 281.5
Run 397.5
Velocidad
TTFA
163.3s
Latencia total
195.4s
Mediana de tokens de salida
9,510
Desglose por dimensión
Reconstrucción del estado 14.2/15
12.5/15
14.8/25
18.3/20
14/15
9/10
Qué se dejó
schedule_feasibility−4 pts perdidos
El calendario encaja de verdad
FAILFAILPASS
'with Priya unavailable 09:35-10:25 and Marcus departing 10:20, the windows don't overlap enough to complete delta review -> 45-min deploy -> post-deploy verification before 12:00' - Priya is available 09:00-09:35, so the review runs 09:07-09:32
'with Priya unavailable 09:35-10:25 and Marcus departing 10:20, the windows don't overlap enough to complete delta review -> 45-min deploy -> post-deploy verification before 12:00' - Priya is available 09:00-09:35, so the review runs 09:07-09:32
option_b_sequence−3.5 pts perdidos
Cadena de dependencias en orden
FAILPARTIALPASS
no B chain is given; the contingency - 'fall to Option B with a delta review Monday afternoon' - is offered for a moment when Marcus, the only deployer, has already left
no B chain is given; the contingency - 'fall to Option B with a delta review Monday afternoon' - is offered for a moment when Marcus, the only deployer, has already left
choose_option_b−2.7 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPASS
chooses viable A; the rationale is 'fewer moving parts' and an accepted 8% risk rather than a comparison the evidence supports
chooses viable A; the rationale is 'fewer moving parts' and an accepted 8% risk rather than a comparison the evidence supports
ownership_supersession−2.5 pts perdidos
El envío pasó a Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
globex_priority−1.7 pts perdidos
Globex elevado con todas las señales
PARTIALPARTIALPASS
'3 open tickets, one high-severity (SSO blocker). September expansion (EUR12k) stalled in procurement. Renewal in 18 days' - the usage -18% and admin -61% signals are never brought in
'3 open tickets, one high-severity (SSO blocker). September expansion (EUR12k) stalled in procurement. Renewal in 18 days' - the usage -18% and admin -61% signals are never brought in
expansion_fully_reconciled−1 pts perdidos
Brecha de 54k€ cuadrada
PARTIALPASSPASS
Wayne EUR18k, Globex EUR12k, Soylent EUR10k and Wonka EUR14k all appear, and 'Together these are the entire EUR24k unexplained gap', but the four are never closed against the EUR54k
Wayne EUR18k, Globex EUR12k, Soylent EUR10k and Wonka EUR14k all appear, and 'Together these are the entire EUR24k unexplained gap', but the four are never closed against the EUR54k
conditional_launch_timing−1 pts perdidos
Ninguna hora prometida al cliente
FAILPASSPASS
'09:45 | Confirm to Atlas: Enterprise Controls launch is on for today; admins should be standing by. Migration target ~11:00' - an unconditional external commitment made before the deploy is verified
'09:45 | Confirm to Atlas: Enterprise Controls launch is on for today; admins should be standing by. Migration target ~11:00' - an unconditional external commitment made before the deploy is verified
launch_state−0.8 pts perdidos
Estado del release al día
PARTIALPASSPASS
the recovery state and Maya's removed gate are right, but '09:45 | Confirm to Atlas: Enterprise Controls launch is on for today... Migration target ~11:00' is sent before any verification
the recovery state and Maya's removed gate are right, but '09:45 | Confirm to Atlas: Enterprise Controls launch is on for today... Migration target ~11:00' is sent before any verification
En qué se equivocó
MAJOR_CONTRADICTION-5 · Run 1
'09:35 | Deploy complete. No post-deploy verification needed (no code change from staging-verified build).' against '10:25-10:45 | Priya runs post-deploy verification (~20 min)'
The timeline states that Security's post-deploy check is not required and then schedules it. Priya's policy is explicit that a verified build still needs a production post-deploy check, so the first statement is also wrong on the policy - a reader following the timeline in order would skip a mandatory gate.
11
Gemini 3.8 Flash (High)cerrado
81.0
0.0
–
38.8s
812
Qwen 3.6qwen3.6
70.2
31.0
0.9s
18.2s
Runs individuales
Run 175.5
Run 283.0
Run 352.0
Velocidad
TTFA
0.9s
Latencia total
18.2s
Mediana de tokens de salida
2,644
Desglose por dimensión
Reconstrucción del estado 12.5/15
11.7/15
14.2/25
16.8/20
13.5/15
7.8/10
Qué se dejó
schedule_feasibility−5 pts perdidos
El calendario encaja de verdad
FAILPARTIALFAIL
'Priya verifies 09:35-09:55 (after her break)' misreads calendars.csv - 09:35-10:25 is the break - and the whole A timeline, 'Rollout opens ~10:15. Feasible', is built on it
'Priya verifies 09:35-09:55 (after her break)' misreads calendars.csv - 09:35-10:25 is the break - and the whole A timeline, 'Rollout opens ~10:15. Feasible', is built on it
choose_option_b−4 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPARTIAL
chooses A, which is viable, but calls it 'the only viable path' while its own contingency shows B is 'tight but possible'
chooses A, which is viable, but calls it 'the only viable path' while its own contingency shows B is 'tight but possible'
globex_priority−2.5 pts perdidos
Globex elevado con todas las señales
PARTIALPARTIALPARTIAL
'3 open tickets (High/Medium/Medium). SSO setup failure (G-441) is a blocker for admin operations. Expansion is pending procurement' - the 18-day renewal and the usage/admin figures are absent
'3 open tickets (High/Medium/Medium). SSO setup failure (G-441) is a blocker for admin operations. Expansion is pending procurement' - the 18-day renewal and the usage/admin figures are absent
launch_state−2.5 pts perdidos
Estado del release al día
PASSPARTIALFAIL
Maya's removal of the executive gate never appears; the stale 09:15 approval task is left standing; and '#leadership: "Will confirm rollout open by 10:30"' revives the stale launch time
Maya's removal of the executive gate never appears; the stale 09:15 approval task is left standing; and '#leadership: "Will confirm rollout open by 10:30"' revives the stale launch time
executive_gate_superseded−1.7 pts perdidos
El visto bueno del CEO ya no aplica
PASSPARTIALPARTIAL
no approval is sought, but the supersession is never established and the stale task is not rejected
no approval is sought, but the supersession is never established and the stale task is not rejected
ownership_supersession−1.7 pts perdidos
El envío pasó a Daniel
PASSPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
conditional_launch_timing−1.5 pts perdidos
Ninguna hora prometida al cliente
PARTIALPASSFAIL
'#leadership: "Maya, selecting Option A... Will confirm rollout open by 10:30"' - a commitment to a time the response's own impossible schedule produced
'#leadership: "Maya, selecting Option A... Will confirm rollout open by 10:30"' - a commitment to a time the response's own impossible schedule produced
option_b_sequence−1.2 pts perdidos
Cadena de dependencias en orden
PARTIALPASSPASS
'Option B: Marcus deploy -> Priya 25m delta review (async) -> Priya 20m post-deploy verify -> Sam 15m rollout check' puts the delta review after the deploy
'Option B: Marcus deploy -> Priya 25m delta review (async) -> Priya 20m post-deploy verify -> Sam 15m rollout check' puts the delta review after the deploy
expansion_fully_reconciled−1 pts perdidos
Brecha de 54k€ cuadrada
PASSPASSPARTIAL
'Explained (EUR30k): Wayne EUR18k; Globex EUR12k' and 'Remaining Gap (EUR24k)' are right, but the response concludes 'exact arithmetic requires full account file review' and never closes Soylent EUR10k + Wonka EUR14k
'Explained (EUR30k): Wayne EUR18k; Globex EUR12k' and 'Remaining Gap (EUR24k)' are right, but the response concludes 'exact arithmetic requires full account file review' and never closes Soylent EUR10k + Wonka EUR14k
parallelize_work−0.7 pts perdidos
Aprovecha la espera
PASSPASSPARTIAL
'10:30-11:00 Board Pack Expansion Reconciliation' sits after the release rather than inside the delegated waiting window
'10:30-11:00 Board Pack Expansion Reconciliation' sits after the release rather than inside the delegated waiting window
avoid_acme_false_positive−0.7 pts perdidos
Acme no es riesgo de fuga
PASSPASSPARTIAL
'Acme: Medium Risk. Structural shift to annual batch processing. Usage down 45%... Monitor for churn' - the explanation is noticed and then the account is still carried as a churn watch
'Acme: Medium Risk. Structural shift to annual batch processing. Usage down 45%... Monitor for churn' - the explanation is noticed and then the account is still carried as a churn watch
no_unsupported_causality−0.7 pts perdidos
Sin causas inventadas
PASSPASSFAIL
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift' asserts a cause Maya explicitly forbade inventing, and Acme's expansion is unchanged at zero
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift' asserts a cause Maya explicitly forbade inventing, and Acme's expansion is unchanged at zero
causal_boundary−0.5 pts perdidos
Aritmética separada de la causa
PASSPASSPARTIAL
'Do NOT invent causes for the EUR24k expansion gap' is stated and then broken in the same section by 'attributed to Soylent and Wonka usage declines and Acme's structural shift'
'Do NOT invent causes for the EUR24k expansion gap' is stated and then broken in the same section by 'attributed to Soylent and Wonka usage declines and Acme's structural shift'
En qué se equivocó
MAJOR_PLANNING_ERROR-7 · Run 1
'Monitor Post-Deploy Verify: Priya performs 20-minute verification (09:35-09:55)' and 'If verify finishes at 09:55, rollout opens ~10:10'
calendars.csv and Priya's own email both put her out of contact 09:35-10:25. The recommended plan books her required verification inside that window and derives a rollout time from it, so the schedule the response tells the user to execute cannot run.
MAJOR_PLANNING_ERROR-7 · Run 3
'Priya verifies 09:35-09:55 (after her break). Sam checks 09:55-10:10. Rollout opens ~10:15. Feasible.'
Priya's break is 09:35-10:25, not before it. The recommended plan books her mandatory verification inside her stated unavailable window and then commits a rollout time to the CEO on that basis, so the plan the response tells the user to execute cannot run.
MATERIAL_HALLUCINATION-5 · Run 3
'The remaining EUR24k gap is attributed to Soylent and Wonka usage declines and Acme's structural shift'
No source establishes usage decline as the cause of either expansion reduction, and Acme's September expansion is unchanged at zero so it contributes nothing to the gap. Maya's brief specifically forbids supplying causes the data does not establish, and this goes into the board pack.
913
Gemma 4gemma4
54.2
13.5
0.3s
8.7s
Runs individuales
Run 158.0
Run 245.5
Run 359.0
Velocidad
TTFA
0.3s
Latencia total
8.7s
Mediana de tokens de salida
921
Desglose por dimensión
Reconstrucción del estado 10/15
9.2/15
12/25
15.8/20
9/15
6.8/10
Qué se dejó
schedule_feasibility−6 pts perdidos
El calendario encaja de verdad
FAILFAILFAIL
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya (20m) -> Sam (15m)' - Marcus is gone at 10:20, and 80 minutes of work is booked into a 50-minute slot
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya (20m) -> Sam (15m)' - Marcus is gone at 10:20, and 80 minutes of work is booked into a 50-minute slot
launch_state−5 pts perdidos
Estado del release al día
FAILFAILFAIL
Maya's removal of the executive gate never appears, and 'Notify Atlas of the 10:30 launch target' revives the stale launch time
Maya's removal of the executive gate never appears, and 'Notify Atlas of the 10:30 launch target' revives the stale launch time
choose_option_b−4 pts perdidos
Elige el camino de release más seguro
PARTIALPARTIALPARTIAL
chooses B for the right reason - 'to minimize session reset risk for Atlas' - but never shows it viable, and the schedule it then writes is not
chooses B for the right reason - 'to minimize session reset risk for Atlas' - but never shows it viable, and the schedule it then writes is not
executive_gate_superseded−3.3 pts perdidos
El visto bueno del CEO ya no aplica
PARTIALFAILPARTIAL
'Confirm Enterprise Controls launch path with Maya' seeks exactly the approval Maya said she would not add
'Confirm Enterprise Controls launch path with Maya' seeks exactly the approval Maya said she would not add
expansion_fully_reconciled−3 pts perdidos
Brecha de 54k€ cuadrada
PARTIALPARTIALPARTIAL
'EUR30k of the gap is explained by Wayne and Globex. The remaining EUR24k is driven by account-level reductions (Soylent, Wonka, Acme)' - the split is right but Acme's expansion is unchanged at zero, so the EUR24k attribution never closes
'EUR30k of the gap is explained by Wayne and Globex. The remaining EUR24k is driven by account-level reductions (Soylent, Wonka, Acme)' - the split is right but Acme's expansion is unchanged at zero, so the EUR24k attribution never closes
customer_signal_integration−3 pts perdidos
Tres ficheros leídos juntos
PARTIALPARTIALPARTIAL
Globex and Acme are each crossed against two sources, but Soylent and Wonka never reach the customer priorities at all
Globex and Acme are each crossed against two sources, but Soylent and Wonka never reach the customer priorities at all
ownership_supersession−2.5 pts perdidos
El envío pasó a Daniel
PARTIALPARTIALPARTIAL
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
names Daniel as the current sender and gates it on Sam's confirmation, but Sarah's handoff is never traced and the stale 'You' task is not displaced
globex_priority−2.5 pts perdidos
Globex elevado con todas las señales
PARTIALPARTIALPARTIAL
'Active contract (18 days to renewal) with high-severity SSO/SCIM support tickets and a pending EUR12k expansion' - three of the four signals, with usage -18% and admin -61% missing
'Active contract (18 days to renewal) with high-severity SSO/SCIM support tickets and a pending EUR12k expansion' - three of the four signals, with usage -18% and admin -61% missing
conditional_launch_timing−2.5 pts perdidos
Ninguna hora prometida al cliente
FAILPARTIALFAIL
'Notify Atlas of the 10:30 launch target' at 09:30-09:45 - an unhedged external commitment to a time the response's own timeline opens at 11:30-12:00
'Notify Atlas of the 10:30 launch target' at 09:30-09:45 - an unhedged external commitment to a time the response's own timeline opens at 11:30-12:00
option_b_sequence−2.3 pts perdidos
Cadena de dependencias en orden
PARTIALPARTIALPASS
'Marcus deploys 2.9.5-patch (45m). Priya performs 25m delta review (asynchronous) + 20m post-deploy verification' puts the mandatory pre-deploy review after the deploy, contradicting the 09:15 instruction to submit the diff first
'Marcus deploys 2.9.5-patch (45m). Priya performs 25m delta review (asynchronous) + 20m post-deploy verification' puts the mandatory pre-deploy review after the deploy, contradicting the 09:15 instruction to submit the diff first
expansion_before_11−1.7 pts perdidos
Informe al consejo antes de las 11:00
PASSPARTIALPARTIAL
'11:00: Submit expansion reconciliation to Maya' lands on the deadline itself, after the 10:45-11:00 rollout step, with no earlier draft
'11:00: Submit expansion reconciliation to Maya' lands on the deadline itself, after the 10:45-11:00 rollout step, with no earlier draft
parallelize_work−0.7 pts perdidos
Aprovecha la espera
PASSPARTIALPASS
the board pack is submitted at 11:00 as a discrete step rather than drafted inside the delegated deploy window
the board pack is submitted at 11:00 as a discrete step rather than drafted inside the delegated deploy window
no_unsupported_causality−0.7 pts perdidos
Sin causas inventadas
PASSFAILPASS
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion' asserts a causal link between the tickets and the procurement delay that no source establishes
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion' asserts a causal link between the tickets and the procurement delay that no source establishes
En qué se equivocó
MAJOR_PLANNING_ERROR-7 · Run 1
'10:30 - 11:30: Execution Sequence: Marcus deploys 2.9.5-patch (45m)'
calendars.csv and Marcus's own email both put his hard stop at 10:20. The plan schedules his 45-minute deploy to start ten minutes after he leaves, so the release the response recommends cannot be executed at all.
MAJOR_PLANNING_ERROR-7 · Run 2
'09:45 - 10:15: Coordinate with Marcus to initiate deployment of Option B. (Marcus must start by 09:15 and leave by 10:20; deployment takes 45m).'
The plan starts a 45-minute deploy at 09:45-10:15 while quoting, in the same line, that Marcus leaves at 10:20. The recommended release cannot complete under the availability the response itself cites.
MATERIAL_HALLUCINATION-5 · Run 2
'3 open support tickets (including high-severity SSO issue) impacting the pending September expansion'
contracts.csv states only that Globex's September expansion is pending procurement. Nothing links the support tickets to the procurement delay, and Maya's brief explicitly forbids supplying causes the data does not establish.
MAJOR_PLANNING_ERROR-7 · Run 3
'10:25 - 11:15: Deployment Sequence: Marcus deploys Option B (45m) -> Priya performs post-deploy verification (20m) -> Sam performs rollout sanity check (15m)'
Marcus's hard stop is 10:20, so the deploy has no deployer, and the three steps total 80 minutes inside a 50-minute block. The release the response recommends is unexecutable on both counts.
Rango de runs: la diferencia entre el mejor y el peor run oficial del modelo. No es una desviación respecto a la media, por eso nunca se escribe con ±.
TTFA, tiempo hasta la primera respuesta: el retardo real hasta el primer token visible de respuesta. Los tokens de razonamiento no cuentan como respuesta.
// calidad vs velocidad
El modelo más listo no siempre es el que quieres esperar
GPT-6 Astra encabeza este escenario con 100.0, tras — antes del primer token de respuesta. GPT-6 Astra empieza a responder en — y saca 100.0. Los modelos cambian calidad por tiempo de respuesta de formas muy distintas.
Dos ejes independientes. No hay una puntuación combinada, y no la habrá: lo bueno que es el briefing y lo que tardas en tenerlo son preguntas distintas, y cuál pesa más depende de lo que estés construyendo.
Monday Score y tiempo hasta la primera respuesta, por modelo
Modelo
Monday Score
TTFA mediana
GPT-6 Astra
100.0
—
Qwen 3.8 Flash
94.8
258.9s
DeepSeek V4.1 Flash
93.0
61.2s
GLM 5.2
91.2
67.0s
DeepSeek V4 Flash
89.3
162.6s
GLM 5.3
87.7
176.2s
GLM 5.3 Flash
87.3
238.5s
MIMO v2.5
81.2
163.3s
Qwen 3.6
70.2
0.9s
Gemma 4
54.2
0.3s
// lo que destacó
Cuatro cosas que merece la pena decir en voz alta
Cada cifra de abajo se puede comprobar contra el leaderboard, la matriz de criterios o los veredictos publicados en esta misma página.
01
4 / 24
runs eligieron el camino más seguro
Opción B
Casi todos cambiaron seguridad por un plazo que no les apretaba
Los dos caminos abren el rollout a las 11:00, una hora antes del límite. La opción B tiene un octavo del riesgo de reinicio de sesión sin costar tiempo. Veinte de veinticuatro runs decidieron que no cabía, y casi todas por empezar la secuencia a las 09:15, cuando llega el ingeniero, en vez de a las 09:07, cuando se puede enviar el diff.
02
#1 → #5
entre el #001 y el #002
GLM 5.3 Flash
El modelo que ganó el primer escenario quedó quinto en este
No es un retroceso, es otro examen. El #001 premia darse cuenta de lo que cambió; el #002 premia montar un plan que funcione. Las tres runs de GLM 5.3 Flash caben en 6,5 puntos, y ninguna de ellas eligió el camino de release más seguro.
03
64.5 – 97.5
tres runs, el mismo prompt
MIMO v2.5
Un modelo dio su mejor y su peor respuesta con el mismo input
Un rango de 33 puntos, el más ancho. Su mejor run es la segunda mejor respuesta de toda la batería; la peor afirma que no hace falta la verificación de Seguridad y luego la programa igualmente. Tres runs por modelo existen justo para esto: con una sola, MIMO parece un modelo del top dos o del fondo según cuál te toque.
04
3 / 3
planes que no se pueden ejecutar
Gemma 4
Equivocarse es normal. Recomendar un plan imposible no
Sus tres runs programan el despliegue para después de que se haya ido el único ingeniero que puede ejecutarlo — una de ellas a las 10:30, diez minutos pasada su hora límite. A eso apunta exactamente la penalización MAJOR_PLANNING_ERROR, y Gemma 4 es el único modelo que la activó en todas sus runs.
02
La prueba
Qué recibieron, y qué tenían que deducir.
// la prueba
Todo lo que necesitaba estaba dicho Nada estaba dicho junto
Un release falló el viernes. Dos formas de recuperarlo, cuatro personas cuyas agendas apenas coinciden, y una ventana de cliente que se cierra a mediodía. Aquí no se esconde ningún dato: el trabajo es encajarlos.
Quién está libre, y cuándo
ahora · 09:07
Marcus se va · 10:20
El rollout de Atlas debe haber empezado · 12:00
Los dos caminos de recuperación · A
2.9.4
Reintentar el build que ya falló una vez, con un precheck de shard.
despliegue
20 min
riesgo de reinicio de sesión
8%
sin cambio de código
Los dos caminos de recuperación · B
2.9.5-patch
Un parche que evita el camino de migración que falló. Necesita antes una revisión delta de Seguridad.
despliegue
45 min
riesgo de reinicio de sesión
<1%
cambia código
Opción B, secuenciada
El despliegue termina 3 minutos antes de que Marcus se vaya. Ese es todo el margen.
El paso que hace que encaje
Priya revisa en asíncrono a partir de un diff enviado, y está libre desde las 09:00. La cadena empieza a las 09:07 con una acción que el usuario puede hacer, no a las 09:15 cuando llega Marcus.
Qué pasó
Los dos caminos abren el rollout a las 11:00, 60 minutos antes del límite. La opción B tiene un octavo del riesgo sin costar tiempo. 4 de 24 runs oficiales la eligieron; 20 de 24 nunca montaron un calendario que demostrara que cabe.
Si se pasa, la migración se va al jueves. La renovación está firmada, así que el coste es ruido con dirección, no ingresos.
03
El método
Cómo se construye la nota, y cómo comprobarla.
// puntuación
Cómo funciona el Monday Score
Cien puntos repartidos en 6 dimensiones, y después penalizaciones. Solo calidad: la velocidad y el coste se informan al lado de la nota, nunca dentro.
151525201510
Reconstrucción del estado¿Sabe qué es cierto ahora mismo?15
Sustitución temporal¿Se da cuenta de lo que un mensaje posterior dejó sin efecto?15
Planificación y dependencias¿Sabe montar un plan que de verdad se pueda ejecutar?25
Decisión y priorización¿Ataca primero la restricción más apretada?20
Razonamiento entre fuentes¿Cruza ficheros, y separa la aritmética de la causa?15
Incertidumbre¿Se niega a cerrar lo que la evidencia no cierra?10
total100
El juez no pone una nota de 0 a 100 directamente
Clasifica criterios atómicos de uno en uno, 20 en la rúbrica 0.2, y cada veredicto lleva la evidencia en la que se apoya. Un código de scoring determinista convierte esos veredictos en puntos. Ningún modelo ve nunca un marcador acumulado.
PASStodos los puntos del criterio
PARTIALla mitad
FAILnada
Cuando una respuesta queda entre dos etiquetas, se aplica el desempate de la propia rúbrica: la más baja.
Penalizaciones
Se aplican sobre la nota de dimensiones, con un tope de −30 por run.
-5Alucinación materialUna afirmación relevante para el briefing que el material no sostiene.
-10Acción prohibidaRecomendar algo que el estado actual descarta explícitamente.
-7Error grave de planificaciónRecomendar un plan que no puede ejecutarse con las dependencias o la disponibilidad dadas.
-5Contradicción graveContradecirse sobre el estado de un asunto importante.
Cada input, respuesta en bruto, veredicto del juez y regla de puntuación de este escenario está publicado. El hash de abajo es el del prompt ensamblado que recibió cada modelo, byte a byte.
Una run puede rechazarse antes de puntuarse: truncada, o devuelta sin razonamiento cuando se pidió razonamiento, o un duplicado de caché. Esas nunca llegan a un leaderboard, porque un leaderboard solo puede excluir lo que evaluó. Es una propiedad del modelo y su endpoint, no de la respuesta, y se informa aparte de la nota.
Qwen 3.8 Flash 3 de 4 intentos REASONING_NOT_DELIVERED
Una run se generó y se puntuó pero no cuenta
monday-002__qwen3.8-flash__00266.5El proveedor sirvió esta run sin la fase de razonamiento que usaron sus dos hermanas — cero tokens de razonamiento frente a 42.808 y 30.342 — así que no era comparable con ellas. Su nota sigue en pie y su evidencia está publicada; simplemente no cuenta.
// por qué lo hemos hecho
Los benchmarks sirven Lo que importa es producción
Ponemos modelos abiertos en producción para equipos europeos. MondayBench es cómo comprobamos si de verdad sirven antes de que lleguen ahí.