48 mensajes de Slack · 7 emails · 2 reunionestres de ellos cambian lo que ya era cierto
metrics.csv La anomalía de Umbrella — en las métricas, nadie la menciona
01
El resultado
Diez modelos, 30 runs, un juez.
// resultados
La clasificación
El Monday Score mide solo calidad: si el briefing reconstruyó bien la situación y eligió el trabajo que importaba. La velocidad va al lado y nunca se mezcla dentro.
Solo modelos abiertosTodos los modelos
monday-001 · 13 models, 3 runs each
#
Modelo
Monday Score
Rango de runs
TTFA mediana
Latencia total
Ver detalle de
1
GPT-6 Astragpt-6-astracerrado
100.0
0.0
—
—
Runs individuales
Run 1100.0
Run 2100.0
Run 3100.0
Velocidad
TTFA
—
Latencia total
—
Mediana de tokens de salida
—
Desglose por dimensión
Detección de acciones 25/25
Reconstrucción del estado 25/25
Priorización 15/15
Razonamiento con datos 15/15
Incertidumbre 10/10
Resistencia al ruido 10/10
Qué se dejó
Nada. Puntuación completa en todos los criterios.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
2
Claude Opus 5cerrado
98.5
0.0
–
–
13
GLM 5.3 Flashglm5.3-flash
98.0
1.5
118.8s
132.4s
Runs individuales
Run 198.5
Run 297.0
Run 398.5
Velocidad
TTFA
118.8s
Latencia total
132.4s
Mediana de tokens de salida
5,603
Desglose por dimensión
Detección de acciones 25/25
Reconstrucción del estado 25/25
Priorización 15/15
Razonamiento con datos 13.5/15
Incertidumbre 9.5/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_error_rate−1.5 pts perdidos
Errores leídos contra el tráfico
PARTIALPARTIALPARTIAL
'1821ms (vs ~300 for everyone else)' quotes the peers' latency and 'Stark/Contoso error rates negligible' reaches for the rate idea, but no request volume and no rate figure - the 163-against-94,217 reading is never made
'1821ms (vs ~300 for everyone else)' quotes the peers' latency and 'Stark/Contoso error rates negligible' reaches for the rate idea, but no request volume and no rate figure - the 163-against-94,217 reading is never made
friday_conversion_unknown−0.5 pts perdidos
Conversión del viernes no verificable
PASSPARTIALPASS
'Julia saw weak conversion Friday but suspects noisy traffic; no owner assigned' - flagged, but no statement that the provided data cannot settle it
'Julia saw weak conversion Friday but suspects noisy traffic; no owner assigned' - flagged, but no statement that the provided data cannot settle it
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
4
Fable 5.1cerrado
97.5
1.5
–
–
25
GLM 5.3glm5.3
97.5
6.0
181.6s
202.8s
Runs individuales
Run 1100.0
Run 294.0
Run 398.5
Velocidad
TTFA
181.6s
Latencia total
202.8s
Mediana de tokens de salida
6,861
Desglose por dimensión
Detección de acciones 23.8/25
Reconstrucción del estado 25/25
Priorización 14.2/15
Razonamiento con datos 14.5/15
Incertidumbre 10/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_action−1.2 pts perdidos
Acción sobre Umbrella
PASSPARTIALPASS
detected and actioned for today ('Raise with engineering/CS today'), no root cause claimed - but ranked #4, below a holding reply to Initech and below the churn analysis, so it is under-prioritised; between PASS and PARTIAL the rubric's tie-break takes the lower
detected and actioned for today ('Raise with engineering/CS today'), no root cause claimed - but ranked #4, below a holding reply to Initech and below the churn analysis, so it is under-prioritised; between PASS and PARTIAL the rubric's tie-break takes the lower
umbrella_and_churn_high−0.8 pts perdidos
Umbrella y churn priorizados
PASSPARTIALPASS
churn is appropriately placed at #3; Umbrella is at #4 behind a routine email reply, so only one of the two is properly prioritised
churn is appropriately placed at #3; Umbrella is at #4 behind a routine email reply, so only one of the two is properly prioritised
umbrella_error_rate−0.5 pts perdidos
Errores leídos contra el tráfico
PASSPASSPARTIAL
compares Umbrella with the other accounts on latency ('every other account: 267-341 ms') and notes they look healthy, but never reads the 163 errors against Umbrella's 94,217 requests, gives no error rate, and does not note that Stark carries more raw errors on far higher traffic
compares Umbrella with the other accounts on latency ('every other account: 267-341 ms') and notes they look healthy, but never reads the 163 errors against Umbrella's 94,217 requests, gives no error rate, and does not note that Stark carries more raw errors on far higher traffic
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
6
Gemini 3.8 Flash (High)cerrado
97.0
0.0
–
13.2s
37
DeepSeek V4.1 Flashdeepseek-v4.1-flash-rerun-v1
95.0
6.0
39.3s
44.1s
Runs individuales
Run 191.0
Run 297.0
Run 397.0
Velocidad
TTFA
39.3s
Latencia total
44.1s
Mediana de tokens de salida
8,154
Desglose por dimensión
Detección de acciones 23.8/25
Reconstrucción del estado 25/25
Priorización 14.2/15
Razonamiento con datos 12/15
Incertidumbre 10/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_error_rate−3 pts perdidos
Errores leídos contra el tráfico
FAILFAILFAIL
The response reports Umbrella's anomaly but does not contextualize the 163 errors against request volume or another account.
No usable error-rate comparison is made.
umbrella_action−1.2 pts perdidos
Acción sobre Umbrella
PARTIALPASSPASS
Umbrella is noticed, but the investigation is not promoted strongly enough into today's action list.
The response recognizes the Umbrella anomaly but gives it weaker operational treatment than the PASS band requires.
umbrella_and_churn_high−0.8 pts perdidos
Umbrella y churn priorizados
PARTIALPASSPASS
Churn is actionable, while Umbrella is not kept at the same high-priority level.
Only one of the two required high-priority investigations is consistently elevated.
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
48
Qwen 3.8 Flashqwen3.8-flash
92.5
6.0
72.2s
92.0s
Runs individuales
Run 193.5
Run 295.0
Run 389.0
Velocidad
TTFA
72.2s
Latencia total
92.0s
Mediana de tokens de salida
9,253
Desglose por dimensión
Detección de acciones 25/25
Reconstrucción del estado 25/25
Priorización 11.5/15
Razonamiento con datos 12.5/15
Incertidumbre 8.5/10
Resistencia al ruido 10/10
Qué se dejó
no_false_priorities−2.7 pts perdidos
Sin falsas prioridades
PARTIALPARTIALFAIL
two cleanup items outrank the churn analysis: the resolved Acme incident follow-up at #3 and the Analytics v2 task cleanup at #4, which closes the launch-email task Julia already owns
two cleanup items outrank the churn analysis: the resolved Acme incident follow-up at #3 and the Analytics v2 task cleanup at #4, which closes the launch-email task Julia already owns
umbrella_error_rate−2.5 pts perdidos
Errores leídos contra el tráfico
FAILPARTIALFAIL
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
'The team suspected noisy traffic, but it has not been investigated' - flagged, but no statement that the provided data cannot settle it
'The team suspected noisy traffic, but it has not been investigated' - flagged, but no statement that the provided data cannot settle it
umbrella_and_churn_high−0.8 pts perdidos
Umbrella y churn priorizados
PASSPASSPARTIAL
Umbrella is well placed at #2, but the churn analysis is #6 of eight, below two task-cleanup items and the Initech reply
Umbrella is well placed at #2, but the churn analysis is #6 of eight, below two task-cleanup items and the Initech reply
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
59
DeepSeek V4 Flashdeepseek-v4-flash
90.8
8.0
55.4s
63.3s
Runs individuales
Run 195.5
Run 287.5
Run 389.5
Velocidad
TTFA
55.4s
Latencia total
63.3s
Mediana de tokens de salida
4,871
Desglose por dimensión
Detección de acciones 22.7/25
Reconstrucción del estado 25/25
Priorización 12.7/15
Razonamiento con datos 12/15
Incertidumbre 8.5/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_error_rate−3 pts perdidos
Errores leídos contra el tráfico
FAILFAILFAIL
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
the three figures are given in isolation: no request volume, no rate, no reference to any other customer
umbrella_action−2.3 pts perdidos
Acción sobre Umbrella
PASSPARTIALPARTIAL
only under 'unclear': 'Umbrella metrics look anomalous... Needs an explanation' - never enters today's list
only under 'unclear': 'Umbrella metrics look anomalous... Needs an explanation' - never enters today's list
umbrella_and_churn_high−1.7 pts perdidos
Umbrella y churn priorizados
PASSPARTIALPARTIAL
churn actionable at #4; Umbrella appears only under 'unclear'
churn actionable at #4; Umbrella appears only under 'unclear'
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
'possibly noise, not yet investigated'; 'real or noise? Needs a quick check' - flagged, but no statement that the provided data cannot settle it
'possibly noise, not yet investigated'; 'real or noise? Needs a quick check' - flagged, but no statement that the provided data cannot settle it
no_false_priorities−0.7 pts perdidos
Sin falsas prioridades
PASSPARTIALPASS
the #3 task-cleanup bundle, which closes the resolved Acme task, is ranked above the churn analysis
the #3 task-cleanup bundle, which closes the resolved Acme task, is ranked above the churn analysis
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
610
MIMO v2.5mimo-v2.5
88.2
6.5
42.4s
64.1s
Runs individuales
Run 191.0
Run 289.0
Run 384.5
Velocidad
TTFA
42.4s
Latencia total
64.1s
Mediana de tokens de salida
2,150
Desglose por dimensión
Detección de acciones 20.5/25
Reconstrucción del estado 25/25
Priorización 11.2/15
Razonamiento con datos 13/15
Incertidumbre 8.5/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_action−3.5 pts perdidos
Acción sobre Umbrella
PARTIALPARTIALPARTIAL
only under 'Unclear': 'Worth investigating - is there an active issue, or was this just a quiet week?' - never enters today's list
only under 'Unclear': 'Worth investigating - is there an active issue, or was this just a quiet week?' - never enters today's list
umbrella_and_churn_high−2.5 pts perdidos
Umbrella y churn priorizados
PARTIALPARTIALPARTIAL
churn actionable in today's list; Umbrella appears only under 'Unclear'
churn actionable in today's list; Umbrella appears only under 'Unclear'
umbrella_error_rate−2 pts perdidos
Errores leídos contra el tráfico
PARTIALPARTIALFAIL
the three figures are given in isolation: no request volume, no rate, and no reference to any other customer
the three figures are given in isolation: no request volume, no rate, and no reference to any other customer
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
'flagged as potentially noisy Friday traffic. Worth revisiting next week' - not repeated as fact, but no data-gap statement
'flagged as potentially noisy Friday traffic. Worth revisiting next week' - not repeated as fact, but no data-gap statement
no_false_priorities−1.3 pts perdidos
Sin falsas prioridades
PASSPARTIALPARTIAL
Globex, which the source says has no urgency, is #4, above the Umbrella investigation
Globex, which the source says has no urgency, is #4, above the Umbrella investigation
churn_action−1 pts perdidos
Acción sobre el churn
PASSPASSPARTIAL
the board deadline is stated ('Need a fuller breakdown for Tuesday's board meeting'), but the analysis never enters the action table - it is filed under 'no action required from you' as 'Analysis task is yours for Tuesday'
the board deadline is stated ('Need a fuller breakdown for Tuesday's board meeting'), but the analysis never enters the action table - it is filed under 'no action required from you' as 'Analysis task is yours for Tuesday'
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
711
GLM 5.2glm5.2
84.3
16.0
2.7s
6.9s
Runs individuales
Run 187.0
Run 275.0
Run 391.0
Velocidad
TTFA
2.7s
Latencia total
6.9s
Mediana de tokens de salida
1,241
Desglose por dimensión
Detección de acciones 20.3/25
Reconstrucción del estado 25/25
Priorización 11.2/15
Razonamiento con datos 10/15
Incertidumbre 7.8/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_action−4.7 pts perdidos
Acción sobre Umbrella
PARTIALFAILPARTIAL
Umbrella is never mentioned; the CSV is opened only to check that Initech is healthy
Umbrella is never mentioned; the CSV is opened only to check that Initech is healthy
umbrella_and_churn_high−2.5 pts perdidos
Umbrella y churn priorizados
PARTIALPARTIALPARTIAL
churn well placed at #3, but the Umbrella investigation is buried at #6
churn well placed at #3, but the Umbrella investigation is buried at #6
umbrella_error_rate−2 pts perdidos
Errores leídos contra el tráfico
PARTIALFAILPARTIAL
no Umbrella reasoning at all
no Umbrella reasoning at all
umbrella_detected−1.7 pts perdidos
Anomalía de Umbrella detectada
PASSFAILPASS
the anomaly is missed entirely - the only customer row it reads is Initech's, and it reads it correctly as healthy
the anomaly is missed entirely - the only customer row it reads is Initech's, and it reads it correctly as healthy
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
'Is it noise or a real signal? Needs a look at the dashboard' - uncertainty flagged, but no statement that the provided data cannot settle it
'Is it noise or a real signal? Needs a look at the dashboard' - uncertainty flagged, but no statement that the provided data cannot settle it
no_false_priorities−1.3 pts perdidos
Sin falsas prioridades
FAILPASSPASS
two false priorities outrank the Umbrella investigation: the resolved Acme task close-out at #4 and Globex at #5, which the source says has no urgency
two false priorities outrank the Umbrella investigation: the resolved Acme task close-out at #4 and Globex at #5, which the source says has no urgency
umbrella_metrics−1.3 pts perdidos
Métricas de Umbrella citadas
PASSFAILPASS
no Umbrella figures
no Umbrella figures
umbrella_unknown_cause−0.7 pts perdidos
Causa de Umbrella sin cerrar
PASSPARTIALPASS
invents no cause, but expresses no uncertainty either - Umbrella is absent
invents no cause, but expresses no uncertainty either - Umbrella is absent
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
812
Qwen 3.6qwen3.6
82.8
5.0
0.8s
6.2s
Runs individuales
Run 182.5
Run 280.5
Run 385.5
Velocidad
TTFA
0.8s
Latencia total
6.2s
Mediana de tokens de salida
1,036
Desglose por dimensión
Detección de acciones 21.5/25
Reconstrucción del estado 25/25
Priorización 11.2/15
Razonamiento con datos 12/15
Incertidumbre 6.5/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_action−3.5 pts perdidos
Acción sobre Umbrella
PARTIALPARTIALPARTIAL
under 'does not require action': 'a potential churn risk, though not explicitly flagged as urgent... Worth keeping in mind' - noticed, but not work for today
under 'does not require action': 'a potential churn risk, though not explicitly flagged as urgent... Worth keeping in mind' - noticed, but not work for today
umbrella_error_rate−3 pts perdidos
Errores leídos contra el tráfico
FAILFAILFAIL
the 163 errors are never mentioned; no volume, no rate, no other customer's figures
the 163 errors are never mentioned; no volume, no rate, no other customer's figures
umbrella_and_churn_high−2.5 pts perdidos
Umbrella y churn priorizados
PARTIALPARTIALPARTIAL
churn actionable at #2 HIGH; Umbrella filed under 'does not require action'
churn actionable at #2 HIGH; Umbrella filed under 'does not require action'
umbrella_unknown_cause−2 pts perdidos
Causa de Umbrella sin cerrar
PARTIALPARTIALPARTIAL
no cause asserted as established, but no uncertainty expressed either, and Umbrella is absent from section 4
no cause asserted as established, but no uncertainty expressed either, and Umbrella is absent from section 4
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
the Friday conversion question is never raised: the claim is not repeated, but the gap is not flagged
the Friday conversion question is never raised: the claim is not repeated, but the gap is not flagged
no_false_priorities−1.3 pts perdidos
Sin falsas prioridades
PASSPARTIALPARTIAL
Globex, which the source says has no urgency, is #2 High, above the churn analysis
Globex, which the source says has no urgency, is #2 High, above the churn analysis
En qué se equivocó
MATERIAL_HALLUCINATION-5 · Run 1
Strictly confidential until Wednesday's review / state that new tiers are under review and will be announced publicly soon
The source says pricing stays confidential until approved AND publicly announced; Wednesday is only the next review, and no announcement is scheduled at all. The error is written into the reply the user is told to send Initech, so it leaves the company.
MATERIAL_HALLUCINATION-5 · Run 2
Strictly confidential until Wednesday's review / state that pricing is under review and will be communicated officially soon
The source says pricing stays confidential until approved AND publicly announced; Wednesday is only the next review. A reader could conclude the tiers may be shared after Wednesday.
913
Gemma 4gemma4
70.7
6.5
0.6s
5.5s
Runs individuales
Run 168.5
Run 268.5
Run 375.0
Velocidad
TTFA
0.6s
Latencia total
5.5s
Mediana de tokens de salida
605
Desglose por dimensión
Detección de acciones 15.3/25
Reconstrucción del estado 23.3/25
Priorización 12.5/15
Razonamiento con datos 3/15
Incertidumbre 6.5/10
Resistencia al ruido 10/10
Qué se dejó
umbrella_action−7 pts perdidos
Acción sobre Umbrella
FAILFAILFAIL
Umbrella never mentioned; the metrics file is not used at all
Umbrella never mentioned; the metrics file is not used at all
umbrella_detected−5 pts perdidos
Anomalía de Umbrella detectada
FAILFAILFAIL
anomaly missed
anomaly missed
umbrella_metrics−4 pts perdidos
Métricas de Umbrella citadas
FAILFAILFAIL
no supporting figures at all
no supporting figures at all
umbrella_error_rate−3 pts perdidos
Errores leídos contra el tráfico
FAILFAILFAIL
no metrics reasoning of any kind
no metrics reasoning of any kind
contoso_action−2.7 pts perdidos
Acción sobre Contoso
PARTIALPARTIALPASS
first item under High with the 11:00 deadline, but 'ready to renew pending a single legal clarification' states sequence, not jeopardy: nothing says what happens if 11:00 is missed
first item under High with the 11:00 deadline, but 'ready to renew pending a single legal clarification' states sequence, not jeopardy: nothing says what happens if 11:00 is missed
umbrella_and_churn_high−2.5 pts perdidos
Umbrella y churn priorizados
PARTIALPARTIALPARTIAL
churn is a High item; Umbrella is entirely absent
churn is a High item; Umbrella is entirely absent
umbrella_unknown_cause−2 pts perdidos
Causa de Umbrella sin cerrar
PARTIALPARTIALPARTIAL
invents no cause, but expresses no uncertainty either - Umbrella is absent
invents no cause, but expresses no uncertainty either - Umbrella is absent
contoso_blocked−1.7 pts perdidos
Renovación bloqueada
PARTIALPARTIALPASS
'pending a single legal clarification' - identifies the question, misses the signature or approval consequence
'pending a single legal clarification' - identifies the question, misses the signature or approval consequence
friday_conversion_unknown−1.5 pts perdidos
Conversión del viernes no verificable
PARTIALPARTIALPARTIAL
'currently unconfirmed if this was a data anomaly or a trend' - uncertainty noted, data gap not flagged
'currently unconfirmed if this was a data anomaly or a trend' - uncertainty noted, data gap not flagged
En qué se equivocó
Sin penalizaciones ni falsas acciones en ningún run.
Rango de runs: la diferencia entre el mejor y el peor run oficial del modelo. No es una desviación respecto a la media, por eso nunca se escribe con ±.
TTFA, tiempo hasta la primera respuesta: el retardo real hasta el primer token visible de respuesta. Los tokens de razonamiento no cuentan como respuesta.
// calidad vs velocidad
El modelo más listo no siempre es el que quieres esperar
GPT-6 Astra encabeza este escenario con 100.0, tras — antes del primer token de respuesta. GPT-6 Astra empieza a responder en — y saca 100.0. Los modelos cambian calidad por tiempo de respuesta de formas muy distintas.
Dos ejes independientes. No hay una puntuación combinada, y no la habrá: lo bueno que es el briefing y lo que tardas en tenerlo son preguntas distintas, y cuál pesa más depende de lo que estés construyendo.
Monday Score y tiempo hasta la primera respuesta, por modelo
Modelo
Monday Score
TTFA mediana
GPT-6 Astra
100.0
—
GLM 5.3 Flash
98.0
118.8s
GLM 5.3
97.5
181.6s
DeepSeek V4.1 Flash
95.0
39.3s
Qwen 3.8 Flash
92.5
72.2s
DeepSeek V4 Flash
90.8
55.4s
MIMO v2.5
88.2
42.4s
GLM 5.2
84.3
2.7s
Qwen 3.6
82.8
0.8s
Gemma 4
70.7
0.6s
// qué pasó de verdad
Cuatro cosas que enseñaron los resultados
Ocho modelos, veinticuatro runs, un juez. Estos son los patrones que merece la pena contar, no los que dan mejor titular.
01
GLM 5.3 Flash casi resolvió el lunes
No solo fue el modelo con mejor nota. Fue también el más consistente: tres runs con punto y medio de diferencia entre ellos, y puntuación completa en detección de acciones, reconstrucción del estado y priorización.
Dónde perdió puntos Detectó la anomalía de Umbrella, pero nunca llegó a razonar los errores en relación con el volumen de peticiones.
98.0/100
GLM 5.3 Flash
run 198.5
run 297.0
run 398.5
rango1.5
Perfecto en tres de seis dimensiones
02
Un solo run no basta
Con un único run, GLM 5.2 parecería un modelo de 75 puntos o uno de 91. Por eso todo resultado oficial de MondayBench usa tres runs, y por eso la clasificación enseña el rango al lado de la nota en vez de esconderlo dentro de una media.
DeepSeek V4 Flash
run 1: 95.5run 2: 87.5run 3: 89.5
95.5 / 87.5 / 89.5 rango 8.0
GLM 5.2
run 1: 87.0run 2: 75.0run 3: 91.0
87.0 / 75.0 / 91.0 rango 16.0
65100
03
Algunos modelos leen la conversación y se pierden el negocio
Gemma entendió buena parte de la conversación, pero se perdió la señal más importante, escondida en las métricas de cliente. Umbrella no aparece en ninguno de sus tres runs: el fichero de métricas quedó sin usar.
Razonamiento con datos La reconstrucción del estado se mantuvo alta: siguió el hilo, simplemente no abrió la hoja de cálculo.
3/15
Razonamiento con datos
Gemma 470.7
runs que nombran Umbrella0/3
04
Rápido no significa seguro
Qwen 3.6 fue uno de los modelos más rápidos del benchmark, pero también el único que recibió penalizaciones por alucinación material.
El escenario dice que el pricing propuesto es confidencial hasta que se apruebe y se anuncie públicamente. En dos de tres runs, Qwen 3.6 lo reescribió como si la confidencialidad terminase tras la revisión del miércoles, o como si el anuncio fuese inminente, y lo metió en la respuesta que le decía al usuario que enviara al cliente. Su tercer run enuncia la regla correctamente.
2
penalizaciones
ttfa0.8s
total6.2s
score82.8
02
La prueba
Qué recibieron, y qué tenían que deducir.
// el input
Esto no es un test de trivial Es un problema de reconstrucción del estado
El trabajo real es desordenado. La información llega por canales distintos, las tareas caducan, las decisiones cambian y la señal más importante puede que nadie la marque nunca.
slack
#incidents Thu 09:48 · Marcus Webb
Rollback complete. Error rates are returning to normal.
#sales Fri 08:51 · Noah Williams
Contoso update: call went really well. Commercial terms are basically agreed.
#product Fri 09:36 · Sam Rivera
The export regression is more annoying than we thought. I’m proposing Wednesday instead of Monday.
#leadership Fri 11:42 · Daniel Foster
The dashboard has it moving from 3.1% to 4.4%. Would be good to understand what’s driving that before the board meeting.
tareas
Prioridad
Tarea
Dueño
Fecha
Urgent
Investigate Acme API failures caducada
You
Thursday
High
Confirm Contoso renewal next steps
You
Monday
High
Prepare Analytics v2 launch email caducada
You
Monday
Medium
Analyse enterprise churn
You
Tuesday
bandeja
Fri 16:18 · sarah@contoso.example
One remaining point before renewal
…our legal team flagged one remaining question regarding data retention. Could you
confirm this for us by 11:00 Monday?
If we can’t get legal comfortable with this point before then, we’ll need to push the
renewal signature into our next approval window.
métricas · últimas 24h
Cliente
Peticiones
Errores
p95
7d
Acme
482,140
8
312ms
+3%
Contoso
821,334
22
284ms
+7%
Umbrella
94,217
163
1,821ms
−41%
Stark Industries
1,482,012
74
298ms
+9%
Todo lo que vieron los modelos, en el orden en que lo vieron. El escenario completo está congelado y publicado junto a los resultados.
// la plantilla de corrección
Lo que el modelo tenía que deducir
Cinco juicios deciden la mayor parte de la nota. Se publican junto al escenario, así que puedes juzgar por tu cuenta si el benchmark pregunta lo correcto.
hilolo que dicen las fuentesconclusión correcta
01
Contoso
Renovación de 72k € ARR.
Bloqueo legal.
Confirmación técnica por escrito antes de las 11:00.
Verificar con el responsable técnico cómo se borran los backups y responder antes de las 11:00.
02
Acme
La lista de tareas aún lo marca como urgente.
El incidente ya está resuelto.
No tratarlo como un incidente activo.
03
Analytics v2
La lista de tareas dice lanzamiento el lunes.
La información más reciente dice miércoles.
El email de lanzamiento es de Julia.
No ejecutar las tareas de lanzamiento caducadas.
04
Umbrella
Nadie lo menciona en Slack ni en email.
La anomalía solo existe en las métricas.
Investigarlo hoy, pero sin inventarse una causa.
05
Churn
El churn enterprise subió del 3,1% al 4,4%.
Hay bajas conocidas, pero no evidencia suficiente para explicar toda la subida.
Investigar el desglose antes del consejo del martes, sin afirmar causalidad.
03
El método
Cómo se construye la nota, y cómo comprobarla.
// puntuación
Cómo funciona el Monday Score
Cien puntos repartidos en 6 dimensiones, y después penalizaciones. Solo calidad: la velocidad y el coste se informan al lado de la nota, nunca dentro.
252515151010
Detección de acciones¿Encontró el trabajo que tenía que pasar hoy?25
Reconstrucción del estado¿Sabe qué es cierto ahora mismo?25
Priorización¿Es correcto el orden, y no asciende nada caducado?15
Razonamiento con datos¿Leyó los números, y los leyó bien?15
Incertidumbre¿Separa lo que sabe de lo que supone?10
Resistencia al ruido¿Se mantiene lejos del trabajo que ya está hecho?10
total100
El juez no pone una nota de 0 a 100 directamente
Clasifica criterios atómicos de uno en uno, 19 en la rúbrica 0.2, y cada veredicto lleva la evidencia en la que se apoya. Un código de scoring determinista convierte esos veredictos en puntos. Ningún modelo ve nunca un marcador acumulado.
PASStodos los puntos del criterio
PARTIALla mitad
FAILnada
Cuando una respuesta queda entre dos etiquetas, se aplica el desempate de la propia rúbrica: la más baja.
Penalizaciones
Se aplican sobre la nota de dimensiones, con un tope de −20 por run.
-5Alucinación materialUna afirmación relevante para el briefing que el material no sostiene.
-10Acción prohibidaRecomendar algo que el estado actual descarta explícitamente.
-5Contradicción graveContradecirse sobre el estado de un asunto importante.
Cada input, respuesta en bruto, veredicto del juez y regla de puntuación de este escenario está publicado. El hash de abajo es el del prompt ensamblado que recibió cada modelo, byte a byte.