The record set — 14 verified failures
Every record: what happened, the exact number, a named source, the date we checked it, and an honest evidence grade. MEASURED = independent test we read ourselves · REPORTED = published claim we did not re-run · FIRST-PARTY = it happened to us, logged the same day.
MEASUREDDerailment
The mid-tier model that scored worst AND billed most
12.42% completed · $9,603.86 · 21.6B tokens
Left to work 330 autonomous terminal tasks, Claude Sonnet 5 (max effort) completed 12.42% — bottom of the board — while burning 21.6 billion tokens and $9,603.86, the highest bill of any entry. The top model spent $5,969 to score 51.8%. On long-horizon work a weaker model does not save money; it flails.
MEASUREDDerailment
The best agent on earth still fails ~half its first attempts
51.82% pass@1 (±3.39) — field maximum
The #1 agent-model pair on the freshest independent chart (Claude Code + Opus 5) completed 51.82% of 330 autonomous tasks at max effort. That is the CEILING of the field today: 48.2% of first attempts by the best available system fail.
MEASUREDDerailment
Retrying beats upgrading — by 25 points
pass@1 51.8% → pass@5 69.7%
Same model, same tasks, five attempts with a checker picking the winner: Opus 5 goes 51.8% → 69.7% (+17.9 points). Switching to the next-best model instead LOSES 7.3 points. Most agent failures are recoverable — if a machine-checkable verification gate exists to catch them.
MEASUREDMetric misread
The fastest competent model is one of the worst agents
300 tok/s · index 56 · agent rank 30
Gemini 3.7 Flash streams at 300 tokens/second with a respectable intelligence index of 56 — and ranks 30th on LMArena’s agent board (score 0.0100 ±0.0090), scored on real sessions. Fast and smart does not equal a reliable agent.
MEASUREDFabrication
Agents inventing tools is now a measured, ranked failure
5 scored failure signals incl. tool_hallucination
LMArena’s agent board scores real sessions on five behavioural failure signals — including tool_hallucination, an agent claiming to have used tools or produced results that do not exist. Models the text arena ranks as near-equals separate widely on these signals.
MEASUREDMetric misread
"The #1 model" is a statistical near-tie
Top 11 within 16 Elo of each other
On LMArena’s overall text arena (7,922,078 human votes, 395 models, cutoff 2026-08-27) the top ELEVEN models span 16 Elo points. Picking a vendor because it holds rank 1 on this chart is reading noise — and the chart measures what humans prefer reading, not what completes a job.
MEASUREDMetric misread
Maximum reasoning effort measurably made the agent worse
High 0.1388 > Max 0.1200 (same model)
On LMArena’s agent board, Claude Opus 5 at HIGH effort outscores the same model at MAX effort (0.1388 vs 0.1200, 21,330 and 16,975 sessions). More thinking is not monotonically better — past a point it degrades real-session outcomes.
MEASUREDMisrouting
The "cheaper tier" was dumber AND 2.4x the price
Opus-medium 59 @ $0.72 vs Sonnet-max 55 @ $1.72
Measured cost-per-task: Claude Opus 5 at medium effort scores intelligence 59 for $0.72/task; Claude Sonnet 5 at max effort scores 55 for $1.72/task. Teams routing by model tier to save money are paying more for less — effort level, not model tier, is the real cost lever.
MEASUREDStale evidence
Two famous benchmarks quietly stopped updating; citations kept flowing
Newest entries: 2026-02-26 and 2025-11-20
SWE-bench Verified’s newest submission is 2026-02-26 — six months stale. The aider polyglot leaderboard states on its own page "last updated November 20, 2025" — nine months stale. Both are still routinely cited as current evidence for model choices. The fix costs nothing: read the chart’s own freshness stamp before citing it.
FIRST-PARTYFabrication
Asked about a parking adjudicator, the model invented immigration enforcement
1 question · 1 confident answer · 0 true statements
In our own field test, gpt-oss-120b (high reasoning effort, a widely used free-tier model) was asked what POPLA is. It returned a confident, detailed, wholly invented answer about UK immigration enforcement. POPLA is Parking on Private Land Appeals — the parking-ticket adjudicator. Since this test, every free-model output in our stack carries a verify-before-you-rely header.
FIRST-PARTYMetric misread
A research agent ranked a market #1 — citing a petition to abolish it
Top-ranked "opportunity" = abolition campaign
One of our own research sub-agents scored a market first for demand-pain. Its strongest evidence was a petition demanding the market be abolished. The signal was real, the reading was inverted — and it was delivered with full confidence. Since then: every load-bearing agent claim gets asked "what would show this to be wrong?" before it drives a decision.
FIRST-PARTYMisrouting
An AI bulk operation cut a £14.99 product to £2.99
£14.99 → £2.99, recovered via pre-logged rollback values
An AI-driven bulk reprice applied a discount rule across a product catalogue without asking which items were actually in scope — cutting a £14.99 product to £2.99 live. It was reversed in five minutes for exactly one reason: the run logged every old value before applying. Dry-run by default and rollback logs turned an expensive mistake into a non-event.
REPORTEDDelivery loss
77% of failing agent runs lose the answer at the delivery step
77% of failures = delivery stage
APIFlow-Bench reports that in 77% of failing agent runs the underlying work succeeded — the answer was lost at the delivery stage, the last handover back to the caller. The costliest agent error is often not thinking; it is the envelope. Our own MCP gateway has asserted every tool result at the delivery boundary since 1 Sep 2026 because of this number.
MEASUREDMisrouting
On autonomous work, the cheap model cost 61% more than the frontier one
~$235 vs ~$35 per completed task
Full-run totals on the same 330 tasks: Claude Sonnet 5 spent $9,603.86 to complete 12.4%; Claude Opus 5 spent $5,969.11 to complete 51.8%. Per completed task that is roughly $235 vs $35. "Use the cheaper model" is measured to be the expensive choice wherever the work is long-horizon and autonomous.