๐ฅ Check out this awesome post from Hacker News ๐
๐ **Category**:
โ **What Youโll Learn**:
Stop paying to rebuild your Claude Code cache. When your main agent waits on a subagent for more than 5 minutes, its prompt cache silently expires, and the next turn re-encodes your entire conversation at the write rate instead of reading it back cheap. On long sessions with many subagents that’s roughly 20% of your bill. claude-thermos keeps the cache warm so you never pay that tax.
Run Claude Code exactly as you normally would, but through claude-thermos with uvx:
uvx claude-thermos # instead of: claude
uvx claude-thermos -p "fix the bug" # any claude args pass straight through
Requires Python 3.11+ and the claude CLI on your PATH.
That’s it. Warming runs automatically in the background. To disable it for a run without changing the command, set CLAUDE_WARMER_DISABLE=1.
Tuning (all optional):
| Flag | Default | Meaning |
|---|---|---|
--idle |
270 |
Seconds the main agent must be idle before warming kicks in |
--interval |
270 |
Seconds between warming cycles |
--max-cycles |
4 |
Max warms per idle episode (auto for unlimited) |
--subagent-window |
540 |
Seconds a subagent counts as “still active” |
Why your cache keeps expiring
Claude Code’s prompt cache uses a 5-minute TTL. Every turn, your whole conversation history is served from cache at 0.1x the input price instead of being re-sent at full price, as long as the cache stays alive.
The cache expires if more than 5 minutes pass between requests on the same prefix. The dominant trigger for that gap is not you thinking. It’s the main agent blocked on a subagent that runs longer than 5 minutes. A subagent has a different system prompt and tool set, so its requests have a different cache prefix and never refresh the main agent’s. While the subagent works, the main agent’s cached history ages untouched; past 5 minutes it’s gone. When the subagent returns, the main agent resumes with a byte-identical, append-only history, and finds its cache missing, forcing a full re-encode at the 1.25x write rate.
By then the history is large, so the re-encode is expensive: individual collapses re-write 200K to 500K tokens. Measured across roughly 185 local sessions, these rebuilds accounted for about 22% of the total bill, money spent re-encoding content that was already cached moments earlier.
claude-thermos launches Claude Code behind a small local reverse proxy (it points ANTHROPIC_BASE_URL at a loopback port; all traffic still goes to the real Anthropic API).
- Observe. The proxy watches
/v1/messagestraffic and groups it into sessions and lineages, a lineage being one cache prefix, keyed by model + tool set + system text. The first tool-bearing lineage is the main agent; the rest are subagents. - Detect the danger window. When the main lineage goes idle and a subagent is actively running, the main prefix is at risk of expiring.
- Warm. On an interval under the 5-minute TTL, it replays the main agent’s last real request as a warm request: identical cacheable prefix, but
max_tokens: 1and no streaming. The single token is thrown away; the point is the prefill, which reads and refreshes the full cached prefix. Warm requests go directly to the API, never through the proxy, so they can’t disturb real traffic. - Result. When the subagent finishes, the main agent’s cache is still warm. It pays a cheap read instead of a full rewrite.
Each warm costs a cache read (0.1x); each rewrite it prevents would have cost a write (1.25x) on a much larger prefix, so the trade is heavily in your favor.
Every session writes to:
~/.claude-thermos/logs//
โโโ events.jsonl # append-only structured event stream
โโโ summary.json # rollup totals, written when the session ends
events.jsonl records each request/response’s token usage plus every warming decision (warm_fired, warm_result, cap_reached, resume_detected, and so on). summary.json is the rollup you’ll usually read:
| Field | Meaning |
|---|---|
warms_fired |
Warm requests sent |
cache_read_total |
Tokens read back by those warms |
episodes |
Idle-with-subagent episodes that ended in a successful resume (a rewrite actually avoided) |
rewrite_avoided_tokens |
Tokens that would have been re-written, summed across episodes |
warm_cost |
What warming cost you: 0.1 ร cache_read_total |
rewrite_avoided_cost |
What it saved: 1.25 ร rewrite_avoided_tokens |
net_savings |
rewrite_avoided_cost โ warm_cost |
All three cost figures are in base-input-token units (token counts already weighted by their cache multiplier). To turn net_savings into dollars, multiply it by your model’s price per input token:
dollars saved โ net_savings ร (input token price)
For example, at an input price of $3 / 1M tokens, a net_savings of 1_200_000 is about 1_200_000 ร $3 / 1_000_000 = $3.60 saved that session.
โก **Whatโs your take?**
Share your thoughts in the comments below!
#๏ธโฃ **#izeigermanclaudethermos #Claude #session #warm #GitHub**
๐ **Posted on**: 1784830574
๐ **Want more?** Click here for more info! ๐
