// ZERO-CONFIG FINOPS & SAFETY GATEWAY FOR AI AGENTS

TokenCap

The open-source proxy that enforces spend limits, intelligent auto-pacing, and loop breakers on OpenAI, Anthropic, and Gemini. Transparently holds rapid bursts and returns an OpenAI-standard 429 before runaway agent loops bill your card.

npx tokencap --open

// 02 · ARCHITECTURE & INTEGRATION

Hard limits on spend.
Seamless drop-in.

Enterprise-grade financial protection, entirely locally. Intercepts runaway agent loops, paces API bursts, and tracks prompt caching discounts—all while masquerading as a standard OpenAI endpoint.

Loop Buster
SHA-256 circuit breaker for agent loops
Cruise Control & Pacing
Enforces time-based sliding-window caps
Prompt Caching
Cross-model discount tracking
Session Limits
Isolated budgets per agent or timespan
OpenAI Standard
Drop-in /v1 API compatibility
Zero Infra
Single binary, local SQLite WAL

NATIVE COMPATIBILITY

Cruise Control holds bursts transparently · Loop Buster trips on flapping · SQLite WAL ledger

// 03 · QUICKSTART

Run locally or deploy
in three steps.

  1. 01

    Start via CLI or Docker

    npx tokencap --open (or docker compose up -d)
  2. 02

    Configure virtual keys & auto-pacing

    "rollingWindowCap": 0.5, "autoPacing": { "enabled": true }, "loopBuster": { "enabled": true }
  3. 03

    Direct Cursor to TokenCap

    Settings → Models → OpenAI API Base: http://localhost:8787/v1