// ZERO-CONFIG FINOPS & SAFETY GATEWAY FOR AI AGENTS
TokenCap
The open-source proxy that enforces spend limits, intelligent auto-pacing, and loop breakers on OpenAI, Anthropic, and Gemini. Transparently holds rapid bursts and returns an OpenAI-standard 429 before runaway agent loops bill your card.
npx tokencap --open// 02 · ARCHITECTURE & INTEGRATION
Hard limits on spend.
Seamless drop-in.
Enterprise-grade financial protection, entirely locally. Intercepts runaway agent loops, paces API bursts, and tracks prompt caching discounts—all while masquerading as a standard OpenAI endpoint.
- Loop Buster
- SHA-256 circuit breaker for agent loops
- Cruise Control & Pacing
- Enforces time-based sliding-window caps
- Prompt Caching
- Cross-model discount tracking
- Session Limits
- Isolated budgets per agent or timespan
- OpenAI Standard
- Drop-in /v1 API compatibility
- Zero Infra
- Single binary, local SQLite WAL
NATIVE COMPATIBILITY
// 03 · QUICKSTART
Run locally or deploy
in three steps.
- 01
Start via CLI or Docker
npx tokencap --open (or docker compose up -d) - 02
Configure virtual keys & auto-pacing
"rollingWindowCap": 0.5, "autoPacing": { "enabled": true }, "loopBuster": { "enabled": true } - 03
Direct Cursor to TokenCap
Settings → Models → OpenAI API Base: http://localhost:8787/v1