pi-warm-cache
Keep the prompt cache warm, keep the bill cold.
Why it exists
Long agent sessions carry huge context - often 100k to 300k tokens in the prompt prefix - and providers cache that prefix so repeat turns stay fast and cheap. But the cache expires the moment you step away for a coffee.
Come back, and the next turn pays a cold-read penalty or rebuilds the whole context from scratch. We wanted those savings to survive the idle gaps.
What it does
- Runs lightweight keepalive probes just before the cache would expire, keeping the prefix warm across Anthropic, OpenAI, Azure OpenAI, and xAI Grok.
- Replays the exact captured payload - it never rebuilds or mutates your context, so warming can never change what the agent sees.
- Spend guardrails and idle cutoffs keep it economical, and it tracks cumulative savings with transparent pricing so you can watch it pay for itself.
- Handles the different cache windows providers use - 5-minute, 30-minute, and 1-hour - automatically.
Under the hood
- A TypeScript Pi extension that hooks the provider request read-only to capture the exact payload, then replays it on a timer - with hard invalidation the instant the prefix drifts, so it never serves a stale cache.
- Process-wide concurrency gating keeps simultaneous warm requests in check, and JSONL diagnostics make every probe auditable.
Why it matters
Running AI at scale is as much about cost control as capability. This is the unglamorous economics - caching, token budgets, spend guardrails - that keeps an AI product's invoice sane, and it is exactly what we bring to client systems that cannot afford a runaway bill.