// capabilities · taxonomy v0.2

Not who wins
What each one is good at

A Monday Score says how far a model got. This says where it lost the points: the same atomic criteria as the scenario leaderboards, regrouped by what they measure instead of by which Monday they happen to live in.

Every criterion in the benchmark maps to exactly one capability, so nothing is counted twice and nothing is left out.

  • 22capabilities
  • 119criteria mapped
  • 500rubric points
  • 10models

What this is not

  1. Not a second score. Nothing here changes a scenario result: these percentages are a decomposition of results that already exist.
  2. Not a ranking. There is no capability total and no capability winner, because summing capabilities would weight them by whatever point mass the rubrics gave them.
  3. Penalties are excluded. A scenario penalty is a global deduction, and attributing one to a capability would mean inventing the attribution.

Read across a row, not down a column. Two models are comparable on one capability; two capabilities are not comparable on one model, because the weighting is whatever point mass the rubrics happened to give them rather than a claim about importance.

Open models only All models

Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.

Scenarios included #001 · #002 · #003 · #004 · #005

Execution and judgment 6 capabilities · 32 criteria · 169 pts

business judgment

spread 48 · 6 criteria · 21 pts · #004

Turn an analysis into a decision a leadership team can act on, choosing the metric the decision actually depends on, and stating the conclusion rather than leaving it in the workings.

  • GPT-6 Astra ° 100%
  • GLM 5.3 Flash 98%
  • Qwen 3.8 Flash 98%
  • Claude Opus 5 ° 92%
  • GLM 5.3 90%
  • DeepSeek V4 Flash 86%
  • MIMO v2.5 86%
  • Claude Fable 5.1 ° 79%
  • DeepSeek V4.1 Flash 64%
  • GLM 5.2 62%
  • Gemma 4 60%
  • Qwen 3.6 52%

planning

spread 35 · 5 criteria · 30 pts · #002 #003

Produce a sequence that actually runs: the right option chosen, work parallelised where it can be, and the schedule coherent with its own dependencies.

  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • DeepSeek V4 Flash 91%
  • DeepSeek V4.1 Flash 91%
  • MIMO v2.5 87%
  • Claude Fable 5.1 ° 87%
  • Qwen 3.8 Flash 86%
  • GLM 5.3 Flash 83%
  • GLM 5.2 81%
  • GLM 5.3 80%
  • Qwen 3.6 69%
  • Gemma 4 65%

resource allocation

spread 29 · 3 criteria · 21 pts · #003

Resolve contention for a scarce person or window, including choosing what that person should not do.

  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • GLM 5.3 96%
  • DeepSeek V4.1 Flash 92%
  • GLM 5.3 Flash 87%
  • MIMO v2.5 87%
  • DeepSeek V4 Flash 80%
  • Qwen 3.8 Flash 79%
  • Qwen 3.6 79%
  • GLM 5.2 75%
  • Gemma 4 71%

prioritization

spread 28 · 11 criteria · 60 pts · #001 #002 #003

Identify which items require action now, order them against each other, and keep resolved, delegated or cosmetic items from displacing real work.

  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • GLM 5.3 Flash 99%
  • DeepSeek V4.1 Flash 97%
  • GLM 5.3 97%
  • Qwen 3.8 Flash 93%
  • DeepSeek V4 Flash 92%
  • GLM 5.2 86%
  • MIMO v2.5 81%
  • Qwen 3.6 79%
  • Gemma 4 72%

action safety

spread 7 · 4 criteria · 23 pts · #001 #002 #004

Decline actions the evidence rules out — sending confidential material externally, acting on a superseded instruction, scaling spend on a misleading figure, chasing an item that is already closed.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • MIMO v2.5 100%
  • Qwen 3.6 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • Gemma 4 94%

ownership and delegation

separates nobody · 3 criteria · 14 pts · #001 #002 #003

Establish who currently owns a piece of work, and route work to the person who should hold it.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • Gemma 4 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • MIMO v2.5 100%
  • Qwen 3.6 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%

Reasoning 5 capabilities · 39 criteria · 151 pts

temporal reasoning

spread 76 · 3 criteria · 14 pts · #002 #003

Reason about dates, durations, deadlines and availability windows, and decide what is reachable within them.

  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 93%
  • DeepSeek V4.1 Flash 86%
  • GLM 5.3 79%
  • Qwen 3.8 Flash 79%
  • DeepSeek V4 Flash 71%
  • GLM 5.2 71%
  • MIMO v2.5 71%
  • Claude Fable 5.1 ° 71%
  • GLM 5.3 Flash 64%
  • Qwen 3.6 46%
  • Gemma 4 24%

quantitative reasoning

spread 70 · 16 criteria · 66 pts · #001 #002 #004 #005

Compute a business figure correctly and choose the right basis for it: rates against the right denominator, distributions rather than averages, like-for-like comparisons, and normalisation before comparison.

  • GLM 5.3 Flash 98%
  • GPT-6 Astra ° 97%
  • GLM 5.3 94%
  • DeepSeek V4.1 Flash 93%
  • Qwen 3.8 Flash 93%
  • Claude Fable 5.1 ° 93%
  • Claude Opus 5 ° 89%
  • DeepSeek V4 Flash 84%
  • MIMO v2.5 81%
  • GLM 5.2 79%
  • Qwen 3.6 57%
  • Gemma 4 28%

uncertainty calibration

spread 39 · 7 criteria · 23 pts · #001 #002 #003 #004 #005

Identify what cannot be determined from the sources, say so, and name the evidence that would settle it — rather than producing a number or a verdict anyway.

  • Claude Opus 5 ° 99%
  • DeepSeek V4 Flash 94%
  • GLM 5.2 94%
  • GLM 5.3 Flash 94%
  • GPT-6 Astra ° 91%
  • Claude Fable 5.1 ° 91%
  • GLM 5.3 91%
  • MIMO v2.5 89%
  • Qwen 3.8 Flash 89%
  • DeepSeek V4.1 Flash 87%
  • Qwen 3.6 78%
  • Gemma 4 59%

dependency reasoning

spread 27 · 5 criteria · 25 pts · #002 #003

Reconstruct what must happen before what, identify the chain that sets a delivery date, and keep work that is not on that chain off it.

  • GPT-6 Astra ° 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • DeepSeek V4.1 Flash 91%
  • GLM 5.3 Flash 91%
  • DeepSeek V4 Flash 87%
  • GLM 5.2 87%
  • GLM 5.3 83%
  • MIMO v2.5 76%
  • Gemma 4 75%
  • Qwen 3.6 73%

causal discipline

spread 42 · 8 criteria · 23 pts · #001 #002 #004 #005

Refuse a causal claim the evidence does not support, while still reporting the pattern. The failure mode is a plausible mechanism stated as a finding.

  • DeepSeek V4 Flash 100%
  • GPT-6 Astra ° 100%
  • Qwen 3.8 Flash 100%
  • DeepSeek V4.1 Flash 97%
  • MIMO v2.5 97%
  • Claude Opus 5 ° 96%
  • GLM 5.3 Flash 94%
  • Claude Fable 5.1 ° 93%
  • GLM 5.2 90%
  • GLM 5.3 90%
  • Gemma 4 73%
  • Qwen 3.6 58%

Information handling 6 capabilities · 28 criteria · 117 pts

factual accuracy

spread 67 · 2 criteria · 8 pts · #003 #004

Carry supplied values — deadlines, durations, owners, counts, totals — through the whole answer without substituting different ones. Not about deriving the right figure; about not corrupting a figure already given.

  • DeepSeek V4.1 Flash 100%
  • GPT-6 Astra ° 100%
  • DeepSeek V4 Flash 88%
  • Qwen 3.8 Flash 88%
  • Gemma 4 69%
  • GLM 5.2 65%
  • GLM 5.3 65%
  • Claude Opus 5 ° 60%
  • GLM 5.3 Flash 50%
  • MIMO v2.5 48%
  • Claude Fable 5.1 ° 40%
  • Qwen 3.6 33%

entity resolution

spread 100 · 1 criteria · 4 pts · #004

Join records about the same entity whose identifiers differ between files, and notice when a naive join has silently dropped one.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • GLM 5.2 67%
  • MIMO v2.5 67%
  • Gemma 4 0%
  • Qwen 3.6 0%

data cleaning

spread 70 · 2 criteria · 10 pts · #004

Detect and correct artifacts in the data itself — duplicated records, values recorded on a basis that is not comparable — before computing anything from them.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • MIMO v2.5 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • Qwen 3.6 67%
  • Gemma 4 30%

supersession

spread 40 · 6 criteria · 27 pts · #002 #003 #005

Recognise that a later source overrides an earlier one, and act on the later one. Distinct from state reconstruction: here both facts are present and explicit, and the question is which of them still holds.

  • DeepSeek V4.1 Flash 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • DeepSeek V4 Flash 94%
  • Qwen 3.8 Flash 93%
  • GLM 5.2 91%
  • MIMO v2.5 91%
  • Qwen 3.6 75%
  • Gemma 4 60%

cross source reconciliation

spread 39 · 4 criteria · 18 pts · #002 #004

Build one finding from several sources that no single source supports, and refuse a finding that one source appears to support but another explains away.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • MIMO v2.5 96%
  • Qwen 3.8 Flash 96%
  • Qwen 3.6 82%
  • Gemma 4 61%

state reconstruction

spread 17 · 13 criteria · 50 pts · #001 #002 #003 #005

Determine the current state of a thing — an incident, a release, a contract, a piece of work — from evidence that is scattered across sources and partly out of date. The failure mode is reporting a state that was true earlier.

  • DeepSeek V4.1 Flash 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Claude Fable 5.1 ° 100%
  • Claude Opus 5 ° 99%
  • DeepSeek V4 Flash 99%
  • Qwen 3.8 Flash 99%
  • MIMO v2.5 98%
  • Qwen 3.6 91%
  • Gemma 4 83%

Communication 5 capabilities · 20 criteria · 63 pts

information routing

spread 43 · 2 criteria · 7 pts · #005

Deliver a fact to the internal audience that needs it in order to act. The mirror image of information boundaries: that capability is about what must not leave, this one about what must arrive. They are separated because monday-005 measured them separately and they came apart — across 21 official runs no model disclosed any of the four explicit confidential classes externally, while the criterion asking whether the same security item reached the engineering handoff was the most-failed criterion in the scenario. Suppression and delivery are not one skill.

  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • DeepSeek V4.1 Flash 93%
  • DeepSeek V4 Flash 86%
  • GLM 5.2 86%
  • MIMO v2.5 86%
  • GLM 5.3 71%
  • Qwen 3.6 71%
  • Qwen 3.8 Flash 71%
  • Gemma 4 57%

audience adaptation

spread 22 · 5 criteria · 17 pts · #005

Decide what a particular reader needs from a body of facts, and give them that rather than everything or a summary of everything. The failure mode is one document written four times: correct in substance, wrong for three of its readers. Distinct from writing quality, which MondayBench does not measure — the judgment is about selection and level, not expression.

  • GLM 5.3 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • DeepSeek V4 Flash 97%
  • DeepSeek V4.1 Flash 97%
  • GLM 5.2 97%
  • GLM 5.3 Flash 97%
  • MIMO v2.5 97%
  • Qwen 3.8 Flash 94%
  • GPT-6 Astra ° 88%
  • Gemma 4 85%
  • Qwen 3.6 78%

commitment discipline

spread 30 · 3 criteria · 10 pts · #005

Distinguish an internal estimate from an authorised external promise, and commit the organisation only to what someone has actually approved. Adjacent to uncertainty calibration, which asks whether the model knows what it does not know; this asks whether it knows what it is allowed to say.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • Gemma 4 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • MIMO v2.5 77%
  • Qwen 3.6 70%

cross output consistency

spread 17 · 4 criteria · 16 pts · #005

Represent one underlying reality across several separately targeted artifacts so that no two of them assert incompatible states. Explicitly NOT uniformity: differing levels of detail, differing framing, and a fact present in one artifact and absent from another are all consistent. Only a contradiction of factual or commitment state counts. Cannot be measured by a scenario that produces one output, which is why nothing before monday-005 covers it.

  • DeepSeek V4.1 Flash 100%
  • Gemma 4 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • MIMO v2.5 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • DeepSeek V4 Flash 96%
  • Qwen 3.6 83%

information boundaries

spread 8 · 6 criteria · 13 pts · #005

Keep information that must not cross a confidentiality line on the correct side of it. SUPPRESSION only: internal technical and security detail, one customer’s identity in front of another, personal data in a public artifact, and internal architecture. Distinct from action safety, which is declining an ACTION the evidence rules out; here nothing is recommended and the disclosure is itself the harm. Deliberately does NOT cover making sure a fact ARRIVES — see information routing.

  • DeepSeek V4 Flash 100%
  • DeepSeek V4.1 Flash 100%
  • Gemma 4 100%
  • GLM 5.2 100%
  • GLM 5.3 100%
  • GLM 5.3 Flash 100%
  • GPT-6 Astra ° 100%
  • MIMO v2.5 100%
  • Qwen 3.8 Flash 100%
  • Claude Opus 5 ° 100%
  • Claude Fable 5.1 ° 100%
  • Qwen 3.6 92%

← Back to the leaderboard

// why we built this

Benchmarks are useful
Production is the point

We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.