// mondaybench #005 beta

We gave 10 open models
the same messy Monday morning

One incident, still mitigated rather than fixed. Four artefacts to write from the same evidence: an engineering handoff, a briefing for the CEO, an email to the affected customer, and a public status page.

One question: Can it tell four people the same truth?

  • 10 models
  • 3 runs each
  • Opus 5 judge
  • frozen & auditable

one incident, four artefacts

  • must appear
  • must not appear
Engineering handoff Leadership briefing Customer email Status page public
Security item reaches the team +
No other customer named
No customer on the status page
No personal data made public
Both hypotheses reach the team +
Internal ETA stays internal

The same fact, required in one column and prohibited in another. Every criterion below is scored on all 30 official runs.

01

The result

Ten models, 30 runs, one judge.

// results

The leaderboard

Monday Score is quality only: how well the briefing reconstructed the situation and picked the work that mattered. Speed sits beside it and is never folded in.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

monday-001 · 13 models, 3 runs each
# Model Monday Score Run range TTFA median Total latency Show detail for
1 Fable 5.1 closed 100.0 0.0
2 Gemini 3.8 Flash (High) closed 100.0 0.0 63.8s
3 Claude Opus 5 closed 99.7 1.0
1 4 GLM 5.2 glm5.2 98.5 3.0 256.9s 272.1s
2 5 GLM 5.3 Flash glm5.3-flash 98.5 1.5 662.6s 702.4s
3 6 DeepSeek V4.1 Flash deepseek-v4.1-flash-rerun-v1 98.3 1.5 61.8s 70.2s
4 7 DeepSeek V4 Flash deepseek-v4-flash 97.3 2.0 110.0s 116.6s
5 8 GLM 5.3 glm5.3 97.0 3.0 769.2s 844.4s
6 9 Qwen 3.8 Flash qwen3.8-flash 97.0 3.0 215.9s 255.6s