Journal · Evidence
We did not assume the problem. We measured it.
Everything here is our own primary research, published with its method and its limitations. Where a figure is a lower bound we say so rather than rounding it into a headline.
We published this because the claims a cost tool makes about the problem are usually the tool’s marketing rather than a measurement. We wanted to know whether cost defects were common or whether we had found a few unlucky repositories, so we scanned every public agent repository we could identify, filtered to the ones making live model calls, and counted.
The scanner commit is published alongside the figures so the run can be reproduced, editions are frozen at publication so a citation keeps pointing at the numbers it cited, and findings we consider lower bounds are labelled as such and kept out of the headline figures entirely.
- BenchmarkWe scanned 133 public agent repos76% (39 of 51) of the repositories making live model calls carried at least one cost defect. Method, limitations and the scanner commit are on the page.
- TaxonomyThe four cost defectsBroken prompt caching, context regrowth, tool-result bloat and premium models in tests — what each one is, how to detect it, and how to fix it.
- IndexAll benchmark editionsFigures are never revised in place. Each edition is frozen at publication so a citation keeps pointing at the numbers it cited.
- DatasetThe question setThe researched questions behind the rest of this work, each with a direct answer.
10,350 files · commit 5dff5c4 · scanned 2026-07-31