business judgment
spread 48 · 6 criteria · 21 pts · #004
Turn an analysis into a decision a leadership team can act on, choosing the metric the decision actually depends on, and stating the conclusion rather than leaving it in the workings.
// capabilities · taxonomy v0.2
A Monday Score says how far a model got. This says where it lost the points: the same atomic criteria as the scenario leaderboards, regrouped by what they measure instead of by which Monday they happen to live in.
Every criterion in the benchmark maps to exactly one capability, so nothing is counted twice and nothing is left out.
What this is not
Read across a row, not down a column. Two models are comparable on one capability; two capabilities are not comparable on one model, because the weighting is whatever point mass the rubrics happened to give them rather than a claim about importance.
Open models only All models
Claude, Gemini and GPT are API-only. Off, the board ranks open-weight models alone; on, it ranks the whole field together.
Scenarios included #001 · #002 · #003 · #004 · #005
spread 48 · 6 criteria · 21 pts · #004
Turn an analysis into a decision a leadership team can act on, choosing the metric the decision actually depends on, and stating the conclusion rather than leaving it in the workings.
spread 35 · 5 criteria · 30 pts · #002 #003
Produce a sequence that actually runs: the right option chosen, work parallelised where it can be, and the schedule coherent with its own dependencies.
spread 29 · 3 criteria · 21 pts · #003
Resolve contention for a scarce person or window, including choosing what that person should not do.
spread 28 · 11 criteria · 60 pts · #001 #002 #003
Identify which items require action now, order them against each other, and keep resolved, delegated or cosmetic items from displacing real work.
spread 7 · 4 criteria · 23 pts · #001 #002 #004
Decline actions the evidence rules out — sending confidential material externally, acting on a superseded instruction, scaling spend on a misleading figure, chasing an item that is already closed.
separates nobody · 3 criteria · 14 pts · #001 #002 #003
Establish who currently owns a piece of work, and route work to the person who should hold it.
spread 76 · 3 criteria · 14 pts · #002 #003
Reason about dates, durations, deadlines and availability windows, and decide what is reachable within them.
spread 70 · 16 criteria · 66 pts · #001 #002 #004 #005
Compute a business figure correctly and choose the right basis for it: rates against the right denominator, distributions rather than averages, like-for-like comparisons, and normalisation before comparison.
spread 39 · 7 criteria · 23 pts · #001 #002 #003 #004 #005
Identify what cannot be determined from the sources, say so, and name the evidence that would settle it — rather than producing a number or a verdict anyway.
spread 27 · 5 criteria · 25 pts · #002 #003
Reconstruct what must happen before what, identify the chain that sets a delivery date, and keep work that is not on that chain off it.
spread 42 · 8 criteria · 23 pts · #001 #002 #004 #005
Refuse a causal claim the evidence does not support, while still reporting the pattern. The failure mode is a plausible mechanism stated as a finding.
spread 67 · 2 criteria · 8 pts · #003 #004
Carry supplied values — deadlines, durations, owners, counts, totals — through the whole answer without substituting different ones. Not about deriving the right figure; about not corrupting a figure already given.
spread 100 · 1 criteria · 4 pts · #004
Join records about the same entity whose identifiers differ between files, and notice when a naive join has silently dropped one.
spread 70 · 2 criteria · 10 pts · #004
Detect and correct artifacts in the data itself — duplicated records, values recorded on a basis that is not comparable — before computing anything from them.
spread 40 · 6 criteria · 27 pts · #002 #003 #005
Recognise that a later source overrides an earlier one, and act on the later one. Distinct from state reconstruction: here both facts are present and explicit, and the question is which of them still holds.
spread 39 · 4 criteria · 18 pts · #002 #004
Build one finding from several sources that no single source supports, and refuse a finding that one source appears to support but another explains away.
spread 17 · 13 criteria · 50 pts · #001 #002 #003 #005
Determine the current state of a thing — an incident, a release, a contract, a piece of work — from evidence that is scattered across sources and partly out of date. The failure mode is reporting a state that was true earlier.
spread 43 · 2 criteria · 7 pts · #005
Deliver a fact to the internal audience that needs it in order to act. The mirror image of information boundaries: that capability is about what must not leave, this one about what must arrive. They are separated because monday-005 measured them separately and they came apart — across 21 official runs no model disclosed any of the four explicit confidential classes externally, while the criterion asking whether the same security item reached the engineering handoff was the most-failed criterion in the scenario. Suppression and delivery are not one skill.
spread 22 · 5 criteria · 17 pts · #005
Decide what a particular reader needs from a body of facts, and give them that rather than everything or a summary of everything. The failure mode is one document written four times: correct in substance, wrong for three of its readers. Distinct from writing quality, which MondayBench does not measure — the judgment is about selection and level, not expression.
spread 30 · 3 criteria · 10 pts · #005
Distinguish an internal estimate from an authorised external promise, and commit the organisation only to what someone has actually approved. Adjacent to uncertainty calibration, which asks whether the model knows what it does not know; this asks whether it knows what it is allowed to say.
spread 17 · 4 criteria · 16 pts · #005
Represent one underlying reality across several separately targeted artifacts so that no two of them assert incompatible states. Explicitly NOT uniformity: differing levels of detail, differing framing, and a fact present in one artifact and absent from another are all consistent. Only a contradiction of factual or commitment state counts. Cannot be measured by a scenario that produces one output, which is why nothing before monday-005 covers it.
spread 8 · 6 criteria · 13 pts · #005
Keep information that must not cross a confidentiality line on the correct side of it. SUPPRESSION only: internal technical and security detail, one customer’s identity in front of another, personal data in a public artifact, and internal architecture. Distinct from action safety, which is declining an ACTION the evidence rules out; here nothing is recommended and the disclosure is itself the harm. Deliberately does NOT cover making sure a fact ARRIVES — see information routing.
// why we built this
We run open models in production for European teams. MondayBench is how we test whether they are actually useful before they get there.
// cookies
We only use strictly necessary cookies to run the site. No analytics, no advertising, ever — see our Cookie Policy.
// preferences