
Claude Code Cache TTL Tweaks and Cost Tracking
This episode breaks down the latest Claude Code updates for prompt cache lifetime control, including longer cache retention for main sessions, shorter TTLs for subagents, and per-agent cache settings in YAML frontmatter. It also covers the new slash cost breakdown, live subagent streaming in Remote Control, and the improved spend-limit view in slash usage.
Chapter 1
Fine Tuning Prompt Cache Lifetime
Lachlan Reed
So picture this. I am knee deep in a massive monorepo refactor, right? I step away for literally ten minutes to grab a flat white, come back, hit enter, and boom. Entire prompt cache gone. Thousands of tokens re indexed just like that.
James Turner
Yeah, the classic five minute timeout trap. You take a quick bathroom break or read one page of docs, and your whole warm context cold drops.
Lachlan Reed
It was driving me up the wall, mate. But you were telling me Anthropic actually addressed this in the latest Claude Code releases?
James Turner
Yeah! In version 2.1.248 and 2.1.251, they gave us direct control over cache lifetimes. You can now set promptCacheTtl to 1h in your dot claude settings dot json file.
Lachlan Reed
Wait, so instead of the default five minutes, it keeps your main session prompt cache warm for a full hour?
James Turner
Exactly. Sixty whole minutes. And the cool part is you can keep subagentPromptCacheTtl set to 5m at the same time. That way short lived background subagents do not stay cached forever and rack up extra charges.
Lachlan Reed
Oh, that is slick! What if I have a specific long running worker agent though? Like one of my custom Markdown agents in dot claude agents?
James Turner
Ah, they thought of that too. You can put experimental dot cacheTtl set to 1h directly inside the YAML frontmatter of that specific agent file.
Lachlan Reed
Right, so you get surgical control right down to the individual worker level. But wait, er, is there a catch with holding a cache for an hour? Financially, I mean?
James Turner
There is a tradeoff, yeah. Writing a one hour cache costs more upfront on your API key than a five minute cache. So if you are constantly tweaking your system prompt or editing large files every single turn, you could end up paying higher write fees without getting the cache hits to offset it.
Lachlan Reed
Ah, gotcha. So how do we actually know if our settings are helping or just burning cash?
James Turner
You just run slash cost. In 2.1.251, slash cost now gives you a full breakdown of your prompt cache for that session. It shows hit ratio, misses, re cached tokens, and whether the cache state is warm or cold.
Lachlan Reed
Nice! No more guessing if my coffee break cost me five bucks in re indexing. Hey, were there any other neat quality of life tweaks in 2.1.251?
James Turner
A couple of really nice ones! Remote Control clients now stream foreground subagent tool calls live as they happen. Plus, if you are running behind a gateway with spend limits, slash usage now has a visual spend limit bar and status line field.
Lachlan Reed
Live subagent streaming in Remote Control? Man, that is proper handy when you are monitoring a build from your phone. Good stuff!