Get 2,500 events tracked for freeSign up now

All posts
Syed Raza5 min readBudgets, Enforcement

A cap on a shared client is not a cap on a run

A dollar cap attached to a model client inherits that object's lifetime, not the run's. Where the counter lives decides when it resets and who shares it.

A spend counter has to live on something, and the object it lives on decides what it bounds. Attach a dollar cap to a model client and you have capped the client: the counter starts when that object is constructed, resets when the object is replaced, and is shared by every agent and every conversation holding a reference to it. The thing you wanted to bound — one run, one conversation, one job — is a span of execution rather than an object, and the two have different lifetimes.

This is a separate question from the identity a budget is keyed to, which is about what a limit can recognise in a request. This one is about plumbing inside your own process: which variable holds the running total, when it was created, and who else can reach it.

The only hook is usually the longest-lived object

An AutoGen discussion opened on 2026-10-04, asking for a hard dollar cap across a whole AgentChat conversation, states the demand and the available mechanism in the same paragraph. The demand:

One team.run(...) can turn into many model calls: every turn of a RoundRobinGroupChat or Swarm, each agent's tool-calling loop, reflection, a termination check that keeps going longer than expected.

and what that leaves you with:

Token usage is visible per call, but nothing stops the whole conversation at a dollar amount, and a loop between two agents can spend a lot before a human notices.

The mechanism the thread reaches for is the model client, because it is the one object every call in the conversation passes through — OpenAIChatCompletionClient "passes default_headers straight to the underlying AsyncOpenAI", so a budget can ride along on every request the client issues.

Now look at the lifetimes. In that design a model client is usually constructed once, handed to several agents, and those agents are assembled into a team that is run many times. The attach point you are given is the longest-lived object in the picture; the unit you are trying to bound is the shortest.

A counter's reset boundary is its object's lifetime

A client constructed at import has a counter that spans the process. If the cap is meant to be per conversation, the second conversation starts from wherever the first one left it, and the ceiling trips on a run whose position in the sequence decides its fate — run seven is refused and run one is not, with identical inputs. The cap is real. It is just measuring the wrong interval.

There are two ways out and both cost something. Reset the counter at the top of each run, and the ceiling now depends on a write that someone has to remember to perform — and that a run killed mid-flight never performs at all, so the next run inherits a counter it cannot see. Or construct a client per conversation, which gives the counter the right window and makes the connection pool, the configuration and the retry policy per-conversation too, when sharing them was the reason the object existed.

The sharing has a second effect. Two conversations running concurrently through one client share one total, so a heavy run trips the other run's ceiling, and afterwards the counter cannot say which conversation consumed what. That is the same shared-instance accounting described in the answers you discard are still billed, arriving from the enforcement side rather than the reporting side.

Reserved spend and settled spend are different numbers

A ceiling checked before the call has to hold headroom for calls already in flight, for the reason set out in a budget check cannot price the call it allows: the cost is not known until the response returns. That hold is a second kind of entry on the same counter, and a LiteLLM pull request merged in the week to 2026-10-06, reworking how its trace-review budgets account for in-flight work, reports what happens when the two are added together:

Temporary budget holds incorrectly reported the monthly limit reached

Its fix names the distinction directly — "Separate settled spending from renewable, five-minute request reservations" — and the two properties in that phrase are both load-bearing. A reservation is separate, so the ceiling binds on money that was actually spent. And it expires, so a run that dies between the hold and the response releases its headroom rather than subtracting it from every later run forever. A hold with no expiry is a counter that only moves in one direction.

The same pull request lists the concurrency case alongside it — "Overlapping runs repeated paid reviews and fragmented the same findings" — which is the other half of a counter keyed to the wrong thing. Work that two overlapping runs both perform is billed twice when the key is the run rather than the unit of work.

The cap you have to arm is the cap you do not have

A Hacker News thread on 2026-10-04, arguing for default hard budget caps, collects the same complaint from outside this stack, about cloud billing rather than tokens:

Even as a hobbyist, even as a most careful and judicious architect and admin, I could not prevent my VPS from incurring costs beyond my control.

The objection there is not that no limit existed. It is that the limit lived somewhere the spending could outrun. A per-conversation cap that must be constructed, passed down and reset for each conversation has the same property in miniature: it protects the code paths where someone remembered to arm it, and the expensive run is rarely one of those.

Five questions to ask of any in-process ceiling

  1. Which object holds the counter, and when was it constructed? Import time, request time or per run — that answer is the window the ceiling actually covers.
  2. What resets it, and what happens if the run dies first? A reset nobody performs is a ceiling that drifts down over the life of the process.
  3. Who else holds a reference? Shared object, shared ceiling: one run's spend refuses another run's calls, and neither the stop nor the attribution can be split apart afterwards.
  4. Is reserved spend separable from settled spend? If not, the limit binds early and reports a breach that has not happened.
  5. Which calls never reach that object? A budget on a client covers the calls made through that client. A sub-agent that constructs its own client, a framework-internal call, a library retry and a memory embedding each go somewhere else, and are outside the ceiling by construction.

How this works in Capsera

Capsera's answer to "where does the counter live" is that it is not on a client object you construct and pass around. capsera.init() patches the provider clients — Anthropic messages, OpenAI chat completions, OpenAI embeddings and Google Gemini, sync and async — so a call is captured whichever client instance issued it, including the ones created inside a framework you did not write.

Identity is held in a ContextVar stack rather than on an object, which means its lifetime is the dynamic extent of the decorated function: @capsera.agent(name="planner") is in scope for everything called underneath it, unwinds when the function returns, and propagates into asyncio tasks and into worker threads that copy context. Nested decorators inherit, so a sub-agent reports as itself under its parent with nothing threaded through call sites, and asyncio.gather does not interleave attribution between concurrent runs.

Budgets sit on that identity. They are scoped per agent, per team or globally, and checked before the provider call, so a blocking budget raises BudgetExceededError and no tokens are spent; they are re-checked at ingest as well. The honest limits: a per-run cost ceiling is a category concept we write about, not a feature that ships — the scopes that exist are agent, team and global — and a call matching no budget proceeds, as do all calls if Capsera itself is unreachable, because a spend tool must not be the reason your application stops.

Two places to look elsewhere. If you want the ceiling in the request path, applied across providers and independent of the application process, LiteLLM is MIT-licensed, free, covers 140+ providers and enforces hard budgets per key, team, organisation and model. If what you need is the content of the conversation that overspent, Langfuse is open source, free and built for trace inspection; Capsera stores hashes and token counts and never prompt or response content, so it cannot answer that question at all. The layering is set out in gateway vs observability vs governance.

Give every agent an identity, a budget, and hard limits.

One line of code. Anthropic, OpenAI, and Google Gemini.

See pricing