Understanding LLM Token Caching
/ 10 min read
Contents
The missing details you need to run production prompt caches: provider minimums, write costs, TTL/eviction, and how the KV cache actually maps to your bill.
What Is an LLM Cache Hit Rate?
An LLM cache hit occurs when the provider reuses part of a prompt it has already processed. The metric that matters is the token hit rate, not a simple hit/miss count. Most requests vary wildly in size, so a request with 100 cached tokens and one with 100,000 cached tokens both count as “a hit” under a naive ratio, even though the second saves far more money.
Example: 100,000 input tokens, 80,000 reused from cache.
How Does Caching Reduce LLM Cost?
LLM billing separates tokens into normal (uncached) input, cached input, and output. Cached input is billed at a lower rate than normal input; output tokens are billed separately and aren’t reduced by caching the input.
Example: normal input at $3/MTok, cached input at $0.30/MTok. 100,000 cached tokens cost $0.03 instead of $0.30, a 90% reduction.
KV Caching Is Also a Memory and Latency Problem
The cost benefit is only part of the story. A prompt cache reuses the model’s processed state during prefill, so a cache hit can also reduce the work required before the model starts generating. That can improve time to first token, especially when the request contains a long context and only a small amount of new information.
This matters most for long-running agents. After twenty minutes of tool calls, the next request may add only a short error message or the result of a file search to a very long conversation. Most of the previous context is unchanged and can be reused.
The tradeoff is memory. KV state consumes GPU memory, and the pressure grows with context length and the number of concurrent users. A rough capacity calculation can become uncomfortable quickly: a model with a 14.5 GB KV footprint per long context would need about 14.5 TB just to hold 1,000 such contexts before accounting for model weights or other runtime memory.
You can move KV state to DRAM or disk to increase capacity, but loading it back costs bandwidth and latency. More cache storage is not automatically faster. The useful measurements are:
- token hit rate;
- prefill latency and time to first token;
- KV memory per request;
- concurrency before eviction or swapping;
- total input, output, storage, and cache-write cost.
The hit rate is a business metric and a systems metric at the same time. A cache can look cheap on the price card and still hurt throughput if the system spends too long moving cached state back into GPU memory.
Effective Cost Depends on More Than the Read Price
The effective input cost includes more than the price of a cache read:
effective cost = uncached input cost+ cached input cost+ cache-write cost+ storage cost+ output costAs an illustrative example, a workload with a $3.00/MTok uncached input rate, a $0.30/MTok cached rate, and a 96.04% cache hit rate has an effective input cost of about $0.407/MTok before output cost. The exact result depends on write frequency, storage duration, and the provider’s billing model.
This framing follows the KV-caching discussion that motivated this section. The number is an illustrative workload calculation, not a universal benchmark.
Prompt Caching and Response Caching Solve Different Problems
Prompt caching is not the same as caching a completed answer. Both avoid repeated work, but they return different things and have different invalidation rules.
| Redis cache | LLM prompt cache |
|---|---|
| Stores application data or complete results | Stores reusable processing of prompt tokens |
| Usually uses an explicit key | May use repeated prompt prefixes or provider-defined keys |
| Returns the cached value directly | The model still generates a new response |
| Application controls TTL and invalidation | Provider often controls expiration and eviction |
| Can store arbitrary objects | Usually caches input processing or model state |
The distinction that trips people up:
A response cache returns an old answer without calling the model, whereas a prompt cache calls the model but reuses processing of the repeated input.
Full-Response Caching Lives in the Application
Full-response caching is managed by the application - Redis, a database, or an object store.
1. Build the complete LLM request.2. Hash the request.3. Check the cache using that hash.4. If hit, return the stored answer.5. If miss, call the LLM.6. Store the answer with a TTL.request_data = { "model": model, "system": system_prompt, "messages": messages, "tools": tools, "temperature": temperature}
cache_key = "llm:response:" + sha256(json.dumps(request_data))
cached_answer = redis.get(cache_key)if cached_answer: metrics.increment("llm_response_cache_hit") return cached_answer
answer = call_llm(request_data)redis.setex(cache_key, 3600, answer)metrics.increment("llm_response_cache_miss")return answerThe cache key must include every input that could change the result: model name/version, system prompt, messages, tool definitions and results, sampling settings, and tenant identity where relevant.
llm:answer:v3:<tenant_id>:<sha256(canonical_request)>Full-response caching eliminates both input-processing and output-generation cost for repeated requests, but only when the exact previous answer is acceptable to return again. It is unsuitable for personalized answers, real-time data, changing tool results, or creative/random output.
Provider Pricing Has Separate Write and Read Costs
Prompt caching has two priced operations, not one:
| Operation | Anthropic | OpenAI | Gemini |
|---|---|---|---|
| Cache write | 1.25x input (5-min), 2x (1-hour) | 1.25x input rate | Standard input rate (still billed to create) |
| Cache read (hit) | 0.1x input (90% off) | 0.1x input (90% off) | Model- and tier-dependent discount |
| Storage | None (TTL based) | None | Model- and date-dependent storage fee |
| Cache lifetime | 5-min or 1-hour tiers | 30-min default and minimum for GPT-5.6+; earlier models vary | TTL you set, subject to provider limits |
- Writing to cache is not always free. The rate depends on the provider, model, and cache mode.
- Reads are discounted, but the exact rate is provider- and model-specific. Do not use the read discount as a proxy for total request savings.
- Gemini’s implicit caching has no storage fee; its explicit caching charges a per-token-per-hour storage fee on top of the discounted read rate.
- Anthropic’s cache cost depends on a TTL tier: a 1-hour write (2x) costs more than a 5-minute write (1.25x).
Worked Example: Anthropic, 5-Minute Cache
Total input: 120,000 tokensCached input: 100,000 tokensUncached input: 20,000 tokensOutput: 2,000 tokensBase input rate: $3 / MTokCached read rate: $0.30 / MTok (0.1x)Write cost (first time, 100k tokens at 1.25x):
100,000 × ($3 × 1.25) / 1,000,000 = $0.375Subsequent read cost:
100,000 × $0.30 / 1,000,000 = $0.03The same 100,000 tokens uncached: $0.30. So after the write, every cache hit saves 90%.
Minimum Cacheable Prefix Length Depends on the Provider
Providers only cache prompts that are long enough. Too short a stable prefix means zero cache benefit.
| Provider | Model | Minimum cacheable prefix |
|---|---|---|
| Anthropic | Current models | Varies by model, from 512 to 4,096 tokens |
| OpenAI | GPT-5.6 and later | 1,024 tokens |
| OpenAI | Earlier models | Varies by model and request settings |
| Gemini | Several Gemini 3.x models | 4,096 tokens |
| Gemini | 2.5 Flash / Pro | 2,048 tokens |
- For some earlier OpenAI models, cached-token reporting is rounded to 128-token increments. GPT-5.6+ has different documented behavior, so check the model-specific rules.
- Anthropic supports both automatic caching and explicit
cache_controlbreakpoints. Use explicit breakpoints when you need control over exactly what is cached. - Gemini’s minimum depends on the model. There is no single minimum shared by all Gemini models.
Rule of thumb: if your repeated stable prefix is shorter than 512-1,024 tokens, prompt caching won’t help regardless of hit rate.
Cache Lifetime and Eviction Are Part of the Cost Model
Cached KV state isn’t stored forever - it consumes GPU memory, the scarcest resource in LLM serving.
- Sliding-window TTL: some providers refresh an entry when it is reused. For example, Anthropic refreshes its 5-minute and 1-hour cache tiers on use, while OpenAI retention varies by model. A steady stream can keep an entry alive, but do not assume that behavior across providers.
- Eviction triggers: age, no reads, or GPU memory pressure.
- Anthropic exposes this as 5-minute vs 1-hour tiers with different write costs.
- Gemini explicit caches let you set the TTL yourself, and you pay storage per hour it persists.
Practical implication: batch requests so frequent calls reuse the same prefix inside the TTL window, rather than letting the cache expire between calls.
Automatic and Explicit Caching Require Different Application Changes
| Provider | Mechanism | What you must do |
|---|---|---|
| OpenAI | Automatic prefix caching above a threshold | Nothing - just order your prompt stably |
| Anthropic | Automatic or explicit caching | Use automatic caching for the default behavior, or mark a breakpoint with cache_control |
| Gemini | Both implicit (auto) and explicit (CachedContent) | Implicit needs nothing; explicit requires creating a cached content object |
“The provider checks internally” is true for OpenAI and Gemini implicit caching. Anthropic also supports automatic caching, while explicit breakpoints let the application control the cache boundary.
{ "model": "claude-opus-5", "messages": [ { "role": "system", "content": [ { "type": "text", "text": "Large stable system instructions...", "cache_control": { "type": "ephemeral" } } ] }, { "role": "user", "content": "What is my refund status?" } ]}Prompt Ordering Determines What the KV Cache Can Reuse
Prompt caching is normally handled inside the model provider or serving system. The application sends an ordinary request; the provider checks whether a reusable prefix is available.
Conceptually, the provider keys on:
model + tokenizer version + prompt prefix + cache namespaceThe cached value is not text or the final answer - it’s an internal representation of the processed prompt (KV state after the transformer has processed it).
Best-cache-hit prompt ordering:
Stable content (first): System instructions, tool schemas, docs, policy, referenceChanging content (last): Current date, user question, user data, latest tool resultIf changing content appears early in the prefix, it breaks the shared prefix and prevents reuse of everything after it. This is the single most common reason for low cache hit rates.
Worked Example: An Order-Status Support Agent
Say you run an e-commerce support agent. Every customer asks some version of “What is the status of my order?” The agent reads stable system instructions and tool schemas, then per-customer order data, then generates a personalized answer.
Bad prompt (dynamic content first):
Current date: 2026-08-25T19:31:00ZRequest ID: req_abc123Customer ID: cust_456
System instructions: You are a support agent...Tool schemas: { "get_order_status": { "order_id": "string" } }Order data for cust_456: { "order_id": "ORD-789", "status": "shipped", "eta": "2026-08-27" }
User question: What is the status of my order?The first three lines change on every request, breaking the shared prefix immediately. Nothing after them, not the system instructions, not the tool schemas, is reusable. Result: 0% token cache hit rate.
Good prompt (stable content first, dynamic content last):
System instructions: You are a support agent...Tool schemas: { "get_order_status": { "order_id": "string" } }Order data template: { "order_id": "<ORDER_ID>", "status": "<STATUS>", "eta": "<ETA>" }User question: What is the status of my order?
Context:Current date: 2026-08-25T19:31:00ZRequest ID: req_abc123Customer ID: cust_456Actual order data: { "order_id": "ORD-789", "status": "shipped", "eta": "2026-08-27" }The first four blocks are byte-identical for every customer asking this question. Only the trailing Context block changes. Every customer, regardless of which order they’re asking about, reuses the same ~500-token cached prefix. Result: ~83% token hit rate (500 cached / 600 total tokens).
A Small Poe Measurement Shows the Direction
I ran the same custom workload through Poe with two prompt layouts:
| Model and layout | Cached input tokens | Total input tokens | Token hit rate | Cost for 3 calls |
|---|---|---|---|---|
| GPT-5.4, cache-breaking layout | 13,440 | 21,159 | 63.5% | $0.0377 |
| GPT-5.4, stable-prefix layout | 15,488 | 21,183 | 73.1% | $0.0323 |
Claude Sonnet 4.6 showed the same direction, but its much longer reasoning output dominated the total cost.
These are exploratory Poe measurements, not a universal provider benchmark. Poe’s bot layer can add its own cached context, so the useful lesson is to measure read tokens, write tokens, output tokens, and total cost separately.
