Agent memory has a shelf life.
Measured on four GPU setups and 7,786 agent sessions: seconds in GPU memory, minutes in RAM, hours on SSD.
One agent, paused. H100 SXM, 128K tokens of context.
- Clock
- Idle 1.0 s
- Memory belongs in
- GPU memory
- Coming back takes
- 0.7 s
Coding agents spend 17.6 times as long waiting as computing: on tools, on tests, on the person at the keyboard. Their memory waits with them.
of an agent's GPU cost on H100 is its memory held in GPU memory while it waits (95% interval 84.0% to 90.8%; 57.9% if every pause is capped at an hour).
lower all-in cost per agent on H100 at 128K tokens when idle memory moves to a RAM tier sized for it, active compute included (2.0× with pauses capped at an hour).
Find the threshold for your hardware.
The order never changes: in all 12 configurations measured, memory belongs in GPU memory first, then RAM, then SSD, then nowhere. What moves is where each threshold falls.
On H100 SXM at 128K tokens, an idle agent's memory is cheapest in GPU memory for 3.7 s, in host RAM until 4.3 min, on local SSD until 10.8 h, and only then is recomputing it cheaper.
Coming back from each tier
- GPU memory0.71 s
- Host RAM1.40 s
- Local SSD8.64 s
- Recompute40.2 s
Time to first token after the pause, median of the measured repetitions.
Holding it for an hour
- GPU memory$1.70
- Host RAM$0.113
- Local SSD$0.0035
- Recomputenothing held
This session's memory at the tier's dated price.
With LMCache's default settings this session comes back in 40.8 s, the same as recomputing: the default buffer is too small to hold it.
Most pauses are short. Most waiting is not.
Every pause in 7,786 recorded coding-agent sessions, 672,232 of them, placed on the ladder. Half last under a second (49.9%): a tool returning, a test finishing. Yet 187 breaks of more than two days, sessions left open overnight or over a weekend, hold 45.8% of all the time agents wait. No single place for memory serves both ends.
- GPU memory
- 67.8%
- of pauses
- Host RAM
- 27.5%
- of pauses
- Local SSD
- 4.5%
- of pauses
- Recompute
- 0.1%
- of pauses
Defaults forget long sessions.
LMCache ships with a 5 GiB CPU buffer per GPU: 20,480 tokens of this model's memory. A longer session does not fit, every request still succeeds, and the agent quietly comes back from scratch. Size the buffer to the sessions it holds, and alarm on the hit counters.
Time to first token when the agent returns.
How it was measured
- Serving
- Qwen3-32B-FP8 on vLLM 0.11 with LMCache 0.3.15, on rented H100 SXM, A100 80 GB, 2x RTX A6000 and B200 machines.
- Tiers
- Each tier's park, hold and resume is timed: KV kept in GPU memory, moved to host RAM or local SSD through LMCache, or dropped and recomputed. Resumes are held to a 2-second time-to-first-token target.
- Thresholds
- A threshold is the pause at which the next tier down becomes cheaper, counting the cost of parking, holding at the tier's dated price, and coming back.
- Agents
- The pauses come from 7,786 recorded Claude Code and Codex sessions in two public traces: TraceLab (SyFI Lab, University of Washington; CC BY 4.0) and cc-traces-weka (SemiAnalysis; Apache-2.0).
- Pauses
- A pause is the gap between an agent's rounds, exactly as the paper measures it: 672,232 of them, 56,393 hours of waiting in all. Shares of time count every pause at its full length; the paper also reports its headline with pauses capped at an hour.
- B200
- Measured on RunPod with FlashInfer, the tuned attention backend for Blackwell. The SSD tier was not measured there, so its ladder has three rungs.
Every number on this page is generated from the measurement records, not typed.