[ GATEWAY SPECIFICATION // REV-2026 ]

Stop runaway agents before they drain your API budget.

A drop-in reverse proxy that catches infinite tool loops and enforces hard spend caps in volatile RAM. Sub-35ms overhead. Zero payload retention.

TRY LIVE IN YOUR TERMINALcurl -I https://circuit-breaker-api.onrender.com/health
Frankfurt (eu-central-1) — 32ms proxy latency — HTTP/2 SSE active
gateway // request traceeu-central-1

12:08:41.018 POST /v1/chat/completions

12:08:41.021 window scan: prompt hash match 1/3

12:08:41.024 spend check: within configured cap

HTTP/2 200 · upstream response passed

Simulated trace · example values

Trust model // implementation details

Shunt — The In-Memory LLM Firewall

Security claims should match the code. Inspect the header flight and verify the public sandbox directly from your terminal.

01 / ZERO-TRUST HEADER FLIGHT

Zero-Trust Header Mode

Your provider credential travels in x-upstream-key and is held in RAM while this request is forwarded.

Zero-trust credential flow through ShuntAn agent sends a request containing an x-upstream-key header through the Shunt proxy, where the key exists in RAM only, then to OpenAI.Agent ClientOpenAI-compatible SDKx-upstream-keyShunt ProxyRAM only · 0-byte disk writekey cleared after request flightupstream requestOpenAIUpstream provider
Zero-Trust Mode: Keys live strictly in volatile memory for the sub-35ms flight of the request. Never written to PostgreSQL. Never written to logs.

02 / ZERO-AUTH TERMINAL SANDBOX

Test the 35ms proxy latency live right now in your terminal.

We don't store your key—it exists only in RAM for this single request. Sandbox calls are capped at $5/hour and $20/day per server process; requests are forwarded to OpenAI using your credential.

curl -X POST https://circuit-breaker-api.onrender.com/v1/chat/completions \
  -H "Authorization: Bearer cb_sandbox_test" \
  -H "x-upstream-key: sk-proj-YOUR_OPENAI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Verify Shunt proxy"}]}'

ARCHITECTURE // REQUEST PATH

Data path and enforcement points

01 / REQUEST FLOW

Volatile-memory request inspection

HTTP/2 · SSE
Shunt request architectureAn agent sends a request through the Shunt RAM hash and spend-cap engine, then to the OpenAI API.Your AgentOpenAI SDK / RESTShunt ProxyRAM Hash & Cap EngineNo prompt payload persistedapi.openai.comUpstream provider<15ms targetHTTP/2 SSEloop detection + hourly / daily spend checks

02 / LOOP STOP RESPONSE

HTTP 200 · finish_reason: stop

{
  "id": "chatcmpl-shunt-1791402483",
  "object": "chat.completion",
  "created": 1791402483,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "[SHUNT ALERT]: Autonomous agent execution halted."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 14,
    "completion_tokens": 0,
    "total_tokens": 14
  },
  "circuit_breaker": {
    "triggered": true,
    "reason": "Prompt repetition threshold exceeded."
  }
}

03 / VOLATILE MEMORY

RAM hash benchmark

<15ms

memory execution target

FingerprintSHA-256
Prompt disk writes0
Payload storageNone

04 / MICRO-DOLLAR ACCOUNTING

Rolling spend ceilings and backpressure

429 BUDGET_LIMIT_BREACHED

Requests are checked against hourly and daily account caps before they are sent upstream. The bars below illustrate cap windows, not live account utilization.

Hourly ceiling$2.00 / hour

CAP THRESHOLD · USAGE NOT SHOWN

Daily ceiling$10.00 / day

CAP THRESHOLD · USAGE NOT SHOWN

Performance figures are engineering targets and depend on deployment, network conditions, and request profile.

INTERACTIVE REQUEST SIMULATION

Loop detection trace

Protocol reference →
gateway-simulation.log
SIMULATED SHUNT

$ agent.run --task "insert order"

The agent repeats the same database write after an unchanged retry.

WINDOW 0/3 · shunt status: POST /v1/chat/completions 200 OK

CLIENT CONFIGURATION

Change the base URL. Keep the SDK.

from openai import OpenAI

client = OpenAI(
    base_url="https://circuit-breaker-api.onrender.com/v1",
    api_key="cb_live_...",
)

FAILURE ANALYSIS // ILLUSTRATIVE SCENARIO

Autopsy of an Overrun

SQL RETRY LOOP

An unhandled SQL error leaves an agent retrying the same tool call at 120 attempts per minute. In this example, the sliding deque recognizes the repeat at attempt three and interrupts the cycle before the remaining requests are sent.

The $680+ / 45-minute cost is a modeled scenario, not measured customer spend. Estimated avoided spend assumes $35 per intercepted incident.

  1. 01 · 00:00

    SQL write fails

    Agent receives retryable error.

  2. 02 · 00:01

    Retry cadence begins

    Unchanged tool plan repeats at 120/min.

  3. 03 · 00:02

    Attempt #3 detected

    Prompt fingerprint crosses repeat threshold.

  4. 04 · 00:02+

    Loop halted

    $680+ illustrative exposure avoided.

Roan de Jager

Built by an Engineer

Built by Roan de Jager (@roandejager) in Norway.

I built Shunt after an autonomous agent got trapped in a recursive tool retry loop and burned through hundreds of dollars in API credits overnight. It's built in Python and FastAPI for sub-35ms raw performance with strict zero-payload retention.