AI Error Observatory · an Axion Labs instrument

Your AI said it with total confidence.
It was wrong.

A made-up fact in a report. An agent that ran all night and delivered nothing. A price change nobody asked for. If an AI failure just cost you an evening — or a customer — you are not doing it wrong. The failure rates are real, they are measured, and knowing them is how you stop paying for them.

48.2%of first attempts FAIL — for the single best AI agent system measured on the freshest independent benchmark
$9,603billed by a "cheaper" mid-tier model to complete 12.4% of its tasks — the worst score on the board
+17.9ptsrecovered by retrying with a verification check — measured to beat switching models

Terminal-Bench 4.0 (board updated 2026-08-29), read 31 Aug 2026. Every number on this page carries its source and check date.

What actually goes wrong — and how often

Free, no email wall, useful even if you never touch our tools. This is what the measured record shows about how AI systems fail in 2026.

It is not just you

The best agent-model pair in the world completes about half its first attempts on realistic autonomous work. Vendors publish capability charts; almost nobody publishes failure taxonomy. That asymmetry is why an AI error feels like a personal mistake. It isn't — it is the operating condition of the whole field right now.

The three findings that save the most money

1) Retrying the same model with a machine check beats upgrading to a "better" model — measured (+17.9 vs −7.3 points). 2) The cheap model can be the expensive choice: on long autonomous work a weaker model flails and out-bills the frontier one. 3) Benchmarks go stale silently — two of the most-cited charts have not updated in 6 and 9 months. Read the chart's own date before you trust it.

Fabrication2 records

The model invents an answer and delivers it with full confidence.

Derailment3 records

On long autonomous runs the agent loops, flails, and burns money without finishing.

Delivery loss1 record

The work succeeded — the answer was lost on the way back to you.

Misrouting3 records

The wrong model or effort level for the job — often costing more to do worse.

Stale evidence1 record

Decisions made on benchmark charts that quietly stopped updating.

Metric misread4 records

The number was real — the conclusion drawn from it was not.

The record set — 14 verified failures

Every record: what happened, the exact number, a named source, the date we checked it, and an honest evidence grade. MEASURED = independent test we read ourselves · REPORTED = published claim we did not re-run · FIRST-PARTY = it happened to us, logged the same day.

MEASUREDDerailment

The mid-tier model that scored worst AND billed most

12.42% completed · $9,603.86 · 21.6B tokens

Left to work 330 autonomous terminal tasks, Claude Sonnet 5 (max effort) completed 12.42% — bottom of the board — while burning 21.6 billion tokens and $9,603.86, the highest bill of any entry. The top model spent $5,969 to score 51.8%. On long-horizon work a weaker model does not save money; it flails.

Source: Terminal-Bench 4.0 leaderboard (Stanford/Harbor/Laude Institute), board updated 2026-08-29 — https://www.tbench.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#sonnet5-tb4-flail
MEASUREDDerailment

The best agent on earth still fails ~half its first attempts

51.82% pass@1 (±3.39) — field maximum

The #1 agent-model pair on the freshest independent chart (Claude Code + Opus 5) completed 51.82% of 330 autonomous tasks at max effort. That is the CEILING of the field today: 48.2% of first attempts by the best available system fail.

Source: Terminal-Bench 4.0 leaderboard, board updated 2026-08-29 — https://www.tbench.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#best-agent-fails-half
MEASUREDDerailment

Retrying beats upgrading — by 25 points

pass@1 51.8% → pass@5 69.7%

Same model, same tasks, five attempts with a checker picking the winner: Opus 5 goes 51.8% → 69.7% (+17.9 points). Switching to the next-best model instead LOSES 7.3 points. Most agent failures are recoverable — if a machine-checkable verification gate exists to catch them.

Source: Terminal-Bench 4.0 pass@k table, board updated 2026-08-29 — https://www.tbench.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#retry-beats-upgrade
MEASUREDMetric misread

The fastest competent model is one of the worst agents

300 tok/s · index 56 · agent rank 30

Gemini 3.7 Flash streams at 300 tokens/second with a respectable intelligence index of 56 — and ranks 30th on LMArena’s agent board (score 0.0100 ±0.0090), scored on real sessions. Fast and smart does not equal a reliable agent.

Source: ArtificialAnalysis model index + LMArena agent leaderboard — https://lmarena.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#speed-is-not-reliability
MEASUREDFabrication

Agents inventing tools is now a measured, ranked failure

5 scored failure signals incl. tool_hallucination

LMArena’s agent board scores real sessions on five behavioural failure signals — including tool_hallucination, an agent claiming to have used tools or produced results that do not exist. Models the text arena ranks as near-equals separate widely on these signals.

Source: LMArena agent leaderboard (scored on real sessions, not votes) — https://lmarena.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#tool-hallucination-is-scored
MEASUREDMetric misread

"The #1 model" is a statistical near-tie

Top 11 within 16 Elo of each other

On LMArena’s overall text arena (7,922,078 human votes, 395 models, cutoff 2026-08-27) the top ELEVEN models span 16 Elo points. Picking a vendor because it holds rank 1 on this chart is reading noise — and the chart measures what humans prefer reading, not what completes a job.

Source: LMArena overall text leaderboard, vote cutoff 2026-08-27 — https://lmarena.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#rank-one-fallacy
MEASUREDMetric misread

Maximum reasoning effort measurably made the agent worse

High 0.1388 > Max 0.1200 (same model)

On LMArena’s agent board, Claude Opus 5 at HIGH effort outscores the same model at MAX effort (0.1388 vs 0.1200, 21,330 and 16,975 sessions). More thinking is not monotonically better — past a point it degrades real-session outcomes.

Source: LMArena agent leaderboard — https://lmarena.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#more-reasoning-worse
MEASUREDMisrouting

The "cheaper tier" was dumber AND 2.4x the price

Opus-medium 59 @ $0.72 vs Sonnet-max 55 @ $1.72

Measured cost-per-task: Claude Opus 5 at medium effort scores intelligence 59 for $0.72/task; Claude Sonnet 5 at max effort scores 55 for $1.72/task. Teams routing by model tier to save money are paying more for less — effort level, not model tier, is the real cost lever.

Source: ArtificialAnalysis Intelligence Index (cost-per-task methodology is AA’s own) — https://artificialanalysis.ai/leaderboards/models
Checked: 2026-08-31 · Record: /api/errors.json#wrong-axis-routing
MEASUREDStale evidence

Two famous benchmarks quietly stopped updating; citations kept flowing

Newest entries: 2026-02-26 and 2025-11-20

SWE-bench Verified’s newest submission is 2026-02-26 — six months stale. The aider polyglot leaderboard states on its own page "last updated November 20, 2025" — nine months stale. Both are still routinely cited as current evidence for model choices. The fix costs nothing: read the chart’s own freshness stamp before citing it.

Source: swebench.com embedded leaderboard data; aider.chat’s own page label — https://www.swebench.com/
Checked: 2026-08-31 · Record: /api/errors.json#stale-benchmark-citation
FIRST-PARTYFabrication

Asked about a parking adjudicator, the model invented immigration enforcement

1 question · 1 confident answer · 0 true statements

In our own field test, gpt-oss-120b (high reasoning effort, a widely used free-tier model) was asked what POPLA is. It returned a confident, detailed, wholly invented answer about UK immigration enforcement. POPLA is Parking on Private Land Appeals — the parking-ticket adjudicator. Since this test, every free-model output in our stack carries a verify-before-you-rely header.

Source: Axion Labs field test, logged same day in scripts/TOOLS.md (axion-readers, public repo) — https://github.com/endrezsoltdios-sketch/axion-readers
Checked: 2026-08-31 · Record: /api/errors.json#popla-fabrication
FIRST-PARTYMetric misread

A research agent ranked a market #1 — citing a petition to abolish it

Top-ranked "opportunity" = abolition campaign

One of our own research sub-agents scored a market first for demand-pain. Its strongest evidence was a petition demanding the market be abolished. The signal was real, the reading was inverted — and it was delivered with full confidence. Since then: every load-bearing agent claim gets asked "what would show this to be wrong?" before it drives a decision.

Source: Axion Labs session log, 31 Jul 2026, written into our operating rules the same day — https://github.com/endrezsoltdios-sketch/axion-readers
Checked: 2026-07-31 · Record: /api/errors.json#petition-inversion
FIRST-PARTYMisrouting

An AI bulk operation cut a £14.99 product to £2.99

£14.99 → £2.99, recovered via pre-logged rollback values

An AI-driven bulk reprice applied a discount rule across a product catalogue without asking which items were actually in scope — cutting a £14.99 product to £2.99 live. It was reversed in five minutes for exactly one reason: the run logged every old value before applying. Dry-run by default and rollback logs turned an expensive mistake into a non-event.

Source: Axion Labs incident, 31 Jul 2026, codified same day as a standing dry-run law — https://github.com/endrezsoltdios-sketch/axion-readers
Checked: 2026-07-31 · Record: /api/errors.json#bulk-reprice-blindness
REPORTEDDelivery loss

77% of failing agent runs lose the answer at the delivery step

77% of failures = delivery stage

APIFlow-Bench reports that in 77% of failing agent runs the underlying work succeeded — the answer was lost at the delivery stage, the last handover back to the caller. The costliest agent error is often not thinking; it is the envelope. Our own MCP gateway has asserted every tool result at the delivery boundary since 1 Sep 2026 because of this number.

Source: APIFlow-Bench, arXiv 2608.29128 (paper claim; we did not re-run the benchmark) — https://arxiv.org/abs/2608.29128
Checked: 2026-09-01 · Record: /api/errors.json#delivery-stage-loss
MEASUREDMisrouting

On autonomous work, the cheap model cost 61% more than the frontier one

~$235 vs ~$35 per completed task

Full-run totals on the same 330 tasks: Claude Sonnet 5 spent $9,603.86 to complete 12.4%; Claude Opus 5 spent $5,969.11 to complete 51.8%. Per completed task that is roughly $235 vs $35. "Use the cheaper model" is measured to be the expensive choice wherever the work is long-horizon and autonomous.

Source: Terminal-Bench 4.0 leaderboard run costs, board updated 2026-08-29 — https://www.tbench.ai/leaderboard
Checked: 2026-08-31 · Record: /api/errors.json#agentic-cost-inversion

Why this observatory exists

We run AI agents in production every day, and we publish verified data for a living. The same discipline we apply to parking-appeal statistics and tax deadlines — claim, source, date, validation — applied to AI's own failures. No vendor money, no leaderboard worship, no invented numbers.

What this is not. This is v1: 14 records (10 measured, 3 first-party), seeded 1 Sep 2026 from our audit of the live benchmark charts and our own production logs. It is not a real-time incident feed, not a model leaderboard, and not advice on which vendor to buy. It grows as we verify — never faster.

Built for agents as much as people

The full record set is a machine door — stable IDs, evidence grades, provenance on every field:

GET https://errors.getaxionlabs.com/api/errors.json

More verified data doors (parking outcomes, tax deadlines, ticket stats) live on our MCP gateway: mcp.getaxionlabs.com — one connection, every door, free.

Keep the receipts coming

New records land here as they are verified. If you want our working notes on building with AI that has to be right — the checks, the failure patterns, the fixes that measured out — the kit is free.

Get the Axion kit — free Report an AI error (with a source)

Error reports need a named source and a date or they cannot become records — that rule is the whole product.