skip to content
Contents

The missing details you need to run production prompt caches: provider minimums, write costs, TTL/eviction, and how the KV cache actually maps to your bill.

What Is an LLM Cache Hit Rate?

An LLM cache hit occurs when the provider reuses part of a prompt it has already processed. The metric that matters is the token hit rate, not a simple hit/miss count. Most requests vary wildly in size, so a request with 100 cached tokens and one with 100,000 cached tokens both count as “a hit” under a naive ratio, even though the second saves far more money.

Token hit rate=cached input tokenstotal input tokens×100\text{Token hit rate} = \frac{\text{cached input tokens}} {\text{total input tokens}} \times 100

Example: 100,000 input tokens, 80,000 reused from cache.

80,000100,000×100=80%\frac{80{,}000}{100{,}000} \times 100 = 80\%

How Does Caching Reduce LLM Cost?

LLM billing separates tokens into normal (uncached) input, cached input, and output. Cached input is billed at a lower rate than normal input; output tokens are billed separately and aren’t reduced by caching the input.

Total cost=(uncached input×normal rate)+(cached input×cached rate)+(output×output rate)\text{Total cost} = (\text{uncached input} \times \text{normal rate}) + (\text{cached input} \times \text{cached rate}) + (\text{output} \times \text{output rate})

Example: normal input at $3/MTok, cached input at $0.30/MTok. 100,000 cached tokens cost $0.03 instead of $0.30, a 90% reduction.

KV Caching Is Also a Memory and Latency Problem

The cost benefit is only part of the story. A prompt cache reuses the model’s processed state during prefill, so a cache hit can also reduce the work required before the model starts generating. That can improve time to first token, especially when the request contains a long context and only a small amount of new information.

This matters most for long-running agents. After twenty minutes of tool calls, the next request may add only a short error message or the result of a file search to a very long conversation. Most of the previous context is unchanged and can be reused.

The tradeoff is memory. KV state consumes GPU memory, and the pressure grows with context length and the number of concurrent users. A rough capacity calculation can become uncomfortable quickly: a model with a 14.5 GB KV footprint per long context would need about 14.5 TB just to hold 1,000 such contexts before accounting for model weights or other runtime memory.

You can move KV state to DRAM or disk to increase capacity, but loading it back costs bandwidth and latency. More cache storage is not automatically faster. The useful measurements are:

  • token hit rate;
  • prefill latency and time to first token;
  • KV memory per request;
  • concurrency before eviction or swapping;
  • total input, output, storage, and cache-write cost.

The hit rate is a business metric and a systems metric at the same time. A cache can look cheap on the price card and still hurt throughput if the system spends too long moving cached state back into GPU memory.

Effective Cost Depends on More Than the Read Price

The effective input cost includes more than the price of a cache read:

effective cost =
uncached input cost
+ cached input cost
+ cache-write cost
+ storage cost
+ output cost

As an illustrative example, a workload with a $3.00/MTok uncached input rate, a $0.30/MTok cached rate, and a 96.04% cache hit rate has an effective input cost of about $0.407/MTok before output cost. The exact result depends on write frequency, storage duration, and the provider’s billing model.

This framing follows the KV-caching discussion that motivated this section. The number is an illustrative workload calculation, not a universal benchmark.

Prompt Caching and Response Caching Solve Different Problems

Prompt caching is not the same as caching a completed answer. Both avoid repeated work, but they return different things and have different invalidation rules.

Redis cacheLLM prompt cache
Stores application data or complete resultsStores reusable processing of prompt tokens
Usually uses an explicit keyMay use repeated prompt prefixes or provider-defined keys
Returns the cached value directlyThe model still generates a new response
Application controls TTL and invalidationProvider often controls expiration and eviction
Can store arbitrary objectsUsually caches input processing or model state

The distinction that trips people up:

A response cache returns an old answer without calling the model, whereas a prompt cache calls the model but reuses processing of the repeated input.

Full-Response Caching Lives in the Application

Full-response caching is managed by the application - Redis, a database, or an object store.

1. Build the complete LLM request.
2. Hash the request.
3. Check the cache using that hash.
4. If hit, return the stored answer.
5. If miss, call the LLM.
6. Store the answer with a TTL.
request_data = {
"model": model,
"system": system_prompt,
"messages": messages,
"tools": tools,
"temperature": temperature
}
cache_key = "llm:response:" + sha256(json.dumps(request_data))
cached_answer = redis.get(cache_key)
if cached_answer:
metrics.increment("llm_response_cache_hit")
return cached_answer
answer = call_llm(request_data)
redis.setex(cache_key, 3600, answer)
metrics.increment("llm_response_cache_miss")
return answer

The cache key must include every input that could change the result: model name/version, system prompt, messages, tool definitions and results, sampling settings, and tenant identity where relevant.

llm:answer:v3:<tenant_id>:<sha256(canonical_request)>

Full-response caching eliminates both input-processing and output-generation cost for repeated requests, but only when the exact previous answer is acceptable to return again. It is unsuitable for personalized answers, real-time data, changing tool results, or creative/random output.

Provider Pricing Has Separate Write and Read Costs

Prompt caching has two priced operations, not one:

OperationAnthropicOpenAIGemini
Cache write1.25x input (5-min), 2x (1-hour)1.25x input rateStandard input rate (still billed to create)
Cache read (hit)0.1x input (90% off)0.1x input (90% off)Model- and tier-dependent discount
StorageNone (TTL based)NoneModel- and date-dependent storage fee
Cache lifetime5-min or 1-hour tiers30-min default and minimum for GPT-5.6+; earlier models varyTTL you set, subject to provider limits
  • Writing to cache is not always free. The rate depends on the provider, model, and cache mode.
  • Reads are discounted, but the exact rate is provider- and model-specific. Do not use the read discount as a proxy for total request savings.
  • Gemini’s implicit caching has no storage fee; its explicit caching charges a per-token-per-hour storage fee on top of the discounted read rate.
  • Anthropic’s cache cost depends on a TTL tier: a 1-hour write (2x) costs more than a 5-minute write (1.25x).

Worked Example: Anthropic, 5-Minute Cache

Total input: 120,000 tokens
Cached input: 100,000 tokens
Uncached input: 20,000 tokens
Output: 2,000 tokens
Base input rate: $3 / MTok
Cached read rate: $0.30 / MTok (0.1x)

Write cost (first time, 100k tokens at 1.25x):

100,000 × ($3 × 1.25) / 1,000,000 = $0.375

Subsequent read cost:

100,000 × $0.30 / 1,000,000 = $0.03

The same 100,000 tokens uncached: $0.30. So after the write, every cache hit saves 90%.

Minimum Cacheable Prefix Length Depends on the Provider

Providers only cache prompts that are long enough. Too short a stable prefix means zero cache benefit.

ProviderModelMinimum cacheable prefix
AnthropicCurrent modelsVaries by model, from 512 to 4,096 tokens
OpenAIGPT-5.6 and later1,024 tokens
OpenAIEarlier modelsVaries by model and request settings
GeminiSeveral Gemini 3.x models4,096 tokens
Gemini2.5 Flash / Pro2,048 tokens
  • For some earlier OpenAI models, cached-token reporting is rounded to 128-token increments. GPT-5.6+ has different documented behavior, so check the model-specific rules.
  • Anthropic supports both automatic caching and explicit cache_control breakpoints. Use explicit breakpoints when you need control over exactly what is cached.
  • Gemini’s minimum depends on the model. There is no single minimum shared by all Gemini models.

Rule of thumb: if your repeated stable prefix is shorter than 512-1,024 tokens, prompt caching won’t help regardless of hit rate.

Cache Lifetime and Eviction Are Part of the Cost Model

Cached KV state isn’t stored forever - it consumes GPU memory, the scarcest resource in LLM serving.

  • Sliding-window TTL: some providers refresh an entry when it is reused. For example, Anthropic refreshes its 5-minute and 1-hour cache tiers on use, while OpenAI retention varies by model. A steady stream can keep an entry alive, but do not assume that behavior across providers.
  • Eviction triggers: age, no reads, or GPU memory pressure.
  • Anthropic exposes this as 5-minute vs 1-hour tiers with different write costs.
  • Gemini explicit caches let you set the TTL yourself, and you pay storage per hour it persists.

Practical implication: batch requests so frequent calls reuse the same prefix inside the TTL window, rather than letting the cache expire between calls.

Automatic and Explicit Caching Require Different Application Changes

ProviderMechanismWhat you must do
OpenAIAutomatic prefix caching above a thresholdNothing - just order your prompt stably
AnthropicAutomatic or explicit cachingUse automatic caching for the default behavior, or mark a breakpoint with cache_control
GeminiBoth implicit (auto) and explicit (CachedContent)Implicit needs nothing; explicit requires creating a cached content object

“The provider checks internally” is true for OpenAI and Gemini implicit caching. Anthropic also supports automatic caching, while explicit breakpoints let the application control the cache boundary.

{
"model": "claude-opus-5",
"messages": [
{
"role": "system",
"content": [
{
"type": "text",
"text": "Large stable system instructions...",
"cache_control": { "type": "ephemeral" }
}
]
},
{ "role": "user", "content": "What is my refund status?" }
]
}

Prompt Ordering Determines What the KV Cache Can Reuse

Prompt caching is normally handled inside the model provider or serving system. The application sends an ordinary request; the provider checks whether a reusable prefix is available.

Conceptually, the provider keys on:

model + tokenizer version + prompt prefix + cache namespace

The cached value is not text or the final answer - it’s an internal representation of the processed prompt (KV state after the transformer has processed it).

Best-cache-hit prompt ordering:

Stable content (first): System instructions, tool schemas, docs, policy, reference
Changing content (last): Current date, user question, user data, latest tool result

If changing content appears early in the prefix, it breaks the shared prefix and prevents reuse of everything after it. This is the single most common reason for low cache hit rates.

Worked Example: An Order-Status Support Agent

Say you run an e-commerce support agent. Every customer asks some version of “What is the status of my order?” The agent reads stable system instructions and tool schemas, then per-customer order data, then generates a personalized answer.

Bad prompt (dynamic content first):

Current date: 2026-08-25T19:31:00Z
Request ID: req_abc123
Customer ID: cust_456
System instructions: You are a support agent...
Tool schemas: { "get_order_status": { "order_id": "string" } }
Order data for cust_456: { "order_id": "ORD-789", "status": "shipped", "eta": "2026-08-27" }
User question: What is the status of my order?

The first three lines change on every request, breaking the shared prefix immediately. Nothing after them, not the system instructions, not the tool schemas, is reusable. Result: 0% token cache hit rate.

Good prompt (stable content first, dynamic content last):

System instructions: You are a support agent...
Tool schemas: { "get_order_status": { "order_id": "string" } }
Order data template: { "order_id": "<ORDER_ID>", "status": "<STATUS>", "eta": "<ETA>" }
User question: What is the status of my order?
Context:
Current date: 2026-08-25T19:31:00Z
Request ID: req_abc123
Customer ID: cust_456
Actual order data: { "order_id": "ORD-789", "status": "shipped", "eta": "2026-08-27" }

The first four blocks are byte-identical for every customer asking this question. Only the trailing Context block changes. Every customer, regardless of which order they’re asking about, reuses the same ~500-token cached prefix. Result: ~83% token hit rate (500 cached / 600 total tokens).

A Small Poe Measurement Shows the Direction

I ran the same custom workload through Poe with two prompt layouts:

Model and layoutCached input tokensTotal input tokensToken hit rateCost for 3 calls
GPT-5.4, cache-breaking layout13,44021,15963.5%$0.0377
GPT-5.4, stable-prefix layout15,48821,18373.1%$0.0323

Claude Sonnet 4.6 showed the same direction, but its much longer reasoning output dominated the total cost.

These are exploratory Poe measurements, not a universal provider benchmark. Poe’s bot layer can add its own cached context, so the useful lesson is to measure read tokens, write tokens, output tokens, and total cost separately.