Primary research
Benchmarks
We scan public agent code to find out how it actually spends money, and publish the method and the limits beside every number. Each edition is dated and frozen; figures are never revised in place.
Editions are frozen at publication. We do not revise figures in place, because a citation that silently changes underneath the person who made it is worse than one that is out of date: the number they checked has to stay the number that is there. A re-run produces a new edition with its own date and its own scanner commit, and the old one keeps its URL.
Every edition publishes the commit of the scanner that produced it, the number of repositories and files covered, and the limitations we know about. Where a finding is a lower bound rather than a rate it is labelled as one and kept out of the headline figures.
2026-07-31
We scanned 133 public AI agent repos. 76% had a cost defect.
Static analysis of 133 open-source Python agent repositories. Of the 51 making direct provider calls, 76% build system prompts in ways that defeat caching, and 1 in 10 run premium models in their test suites.
Read the full method and data
Next edition: a re-run at roughly 500 repositories once framework-level detection ships, closing the blind spot around frameworks that route provider calls internally.