Capsera

Pricing you can budget for

Start free with 2,500 events a month. Pay when it is load-bearing.

Free

$0/mo

For trying it on a side project.

2,500 events per month

  • 2,500 events per month
  • Live usage dashboard
  • Core cost analytics
  • Anthropic, OpenAI + Gemini tracking
  • Community support
Start free

Pro

Recommended

$149/mo

For a team shipping one product.

Unlimited events

  • Unlimited events
  • Budgets & alerts per agent, team, or org
  • Smart model routing engine
  • Prompt analysis & caching insights
  • Cost forecasting & anomaly detection
  • Priority support
Start Pro

Enterprise

Contact us

For regulated and high-volume estates.

Unlimited events

  • Everything in Pro
  • Model cost optimization engine
  • AI optimizer agent
  • Team management & roles
  • Custom integrations
  • Private deployment options
  • Custom contract and SLA
  • Named support engineer
  • Onboarding and migration help

Works with the 13 providers your agents already call.

  • Anthropic
  • OpenAI
  • Gemini
  • DeepSeek
  • Mistral
  • Bedrock
  • Groq
  • Kimi

From our repo scan

  • Public repos scanned133
  • Files read10,350
  • Making live LLM calls51
  • Agents discovered293
  • Prompts defeating caching347
  • Repos with a cost defect76%

Compare plans

Enterprise

Tracking

  • Events per month2,500UnlimitedUnlimited
  • Live usage dashboard
  • Core cost analytics
  • Anthropic, OpenAI and Gemini tracking

Control

  • Budgets and alerts per agent, team or orgNot included
  • Smart model routing engineNot included
  • Prompt analysis and caching insightsNot included
  • Cost forecasting and anomaly detectionNot included

Optimization

  • Model cost optimization engineNot includedNot included
  • AI optimizer agentNot includedNot included

Team

  • Team management and rolesNot includedNot included
  • Custom integrationsNot includedNot included

Deployment and contract

  • Private deployment optionsNot includedNot included
  • Custom contract and SLANot includedNot included

Support

  • SupportCommunityPriorityNamed engineer
  • Onboarding and migration helpNot includedNot included

Frequently asked questions

Will the SDK slow down my LLM calls?

Tracking never does: events go into a non-blocking in-memory queue and a background thread flushes them every 500ms. Budget enforcement adds one lightweight pre-call check, and it is optional. Turn it off and Capsera is fully out of your call path. Either way, if the backend is ever unreachable, everything fails open and your calls proceed untouched.

You're not a proxy. How can you enforce a hard limit?

Enforcement lives in the SDK, inside your own process. Before each provider call it checks the budgets that match the calling agent; a budget with a block action stops the call before any provider spend, while throttle and downgrade reshape it. If Capsera is ever unreachable the check fails open, so you lose enforcement for that moment, never availability.

Which providers are supported?

Thirteen, across three kinds. Model APIs: Anthropic, OpenAI, Google, Mistral, Cohere, DeepSeek and Kimi. Inference platforms: Groq, Together, Fireworks and Cerebras. Cloud model services: Vertex AI and Bedrock. Sync and async clients both. The SDK patches the provider libraries automatically at init, so there is no proxy, no gateway, and no change to your existing call sites.

How do I put every agent on a budget?

Name each agent where it makes its calls, with a decorator or a context manager, then give it a budget envelope with an action at the limit: alert, throttle, or block. Because identity is inherited through the call context, an agent that spawns sub-agents covers them too, so putting every agent on a budget does not mean enumerating them by hand. Budgets can also be scoped to a team or the whole organisation, and the scopes compose.

Can I set a budget for one agent instead of a whole API key?

Yes, and that is the main reason to use Capsera rather than a provider spend limit. A key-level cap stops every agent sharing that credential when one misbehaves; a per-agent envelope stops only the offender and leaves the rest running. Budgets scope to a single agent, a team, or the organisation, over daily, weekly, or monthly periods with calendar-aware resets.

Does this work with LangGraph and CrewAI?

Capsera instruments the provider clients rather than the framework, so calls made through LangGraph, CrewAI, LangChain, AutoGen, LlamaIndex or plain SDK code are all captured, and helpers attribute spend to LangGraph nodes and CrewAI agents and tasks. One honest caveat: CrewAI routes some calls through litellm internally, so direct client patching can miss those. Framework-level interception that closes this gap is the next scanner change.

Is my prompt content stored?

Never. Prompt analysis extracts structural metadata only: token estimates, message counts, cache usage, context window utilization. The text of your prompts and completions stays in your infrastructure.

How do budgets and alerts work?

Budgets scope to a single agent, a team, or your whole organization, over daily, weekly, or monthly periods with calendar-aware resets. Each budget picks its action at the limit, either alert, throttle, or block, and you get an alert at your threshold (80% by default) plus an exceeded alert at 100%.

How accurate is the cost tracking?

Costs are computed from exact per-model pricing tables using decimal arithmetic, based on the real input, output, and cache token counts returned by each API response, not estimates.

What happens when I hit the free tier limit?

Ingestion returns a clear limit response with your current usage and reset date, and the SDK stops sending events until the month rolls over. Hitting the plan limit never blocks your LLM calls. Only tracking pauses.

Built for every team that ships agents.
Six industries, six roles, one bill.

Know exactly where every token goes.

Get started