Overview

The provider reuses the computed prefix of an earlier request when a new request starts with identical content (definition in prompt-cache). On Claude, a cache read costs 0.1x the base input price on most models (0.05x on Opus 5.5, 0.025x on Fable 5.1), a 5-minute write costs 1.25x, and a 1-hour write costs 2x. Hit rate is a layout problem: stable content first, mutable content last, deterministic order in between.

Keep the longest cacheable prefix stable

The cache key is the prefix, so any change early in the prompt invalidates everything after it. Claude renders tools, then system, then messages:

1. Tool definitions            (stable)
2. System message and policy   (stable)
3. Few-shot examples           (stable)
4. Retrieved context           (per request, deterministic order)
5. Conversation history        (append-only)
6. Current user query          (per request)

On Claude, a change to tool definitions invalidates the whole cache, toggling web search or citations invalidates system and messages, and changing tool_choice, images, thinking parameters, or output_config.effort invalidates the messages portion.

Put per-request content at the end

A timestamp, session ID, or “today is 2026-10-01” at the top of the prompt destroys the cache. Move it below the cached region, into the user message.

Bad:  System: You are a helpful assistant. Today is 2026-10-01.
Good: System: You are a helpful assistant.
      User: What is the weather? (Today is 2026-10-01.)

On Claude models that support mid-conversation system messages (Fable 5.1, Fable 5, Opus 5.5, Opus 5, Opus 4.8, Sonnet 5.5), append a {"role": "system"} message to change instructions mid-session without editing the cached system field.

Sort retrieved chunks deterministically

Sort retrieved chunks by a stable key (chunk ID, document timestamp) instead of relevance score, so a recurring chunk set produces identical bytes. This only pays off when chunk sets recur. Where relevance order measurably improves quality, make the trade explicitly and confirm it with an eval. See rag.

Set provider cache controls explicitly

  • Claude: add a top-level cache_control: {"type": "ephemeral"} for automatic caching (the recommended start), or place it on individual blocks for up to 4 explicit breakpoints. The default TTL is 5 minutes; "ttl": "1h" costs more to write. The minimum cacheable prefix is model-dependent, from 512 tokens (Fable 5.x, Opus 5.x, Sonnet 5.5) to 4,096 (Haiku 4.5, Opus 4.6); a shorter prefix silently does not cache. A cache entry exists only after the first response begins, so send parallel requests after the first returns.
  • OpenAI: caching is automatic for prompts of 1,024 tokens or more on GPT-5.6 and later; use a stable prompt_cache_key to group related requests. GPT-5.6 and later also accept explicit breakpoints through prompt_cache_options, with a 1.25x cache-write charge; earlier models use prompt_cache_retention.
  • Gemini: implicit caching is on by default for 2.5 and newer models, with a per-model minimum token count; create explicit cached-content objects for very long stable contexts.

Check current prices and minimums in cost-control and the provider docs before budgeting.

Hold effort and model constant within a cached session

Prompt caches are per model, so a model upgrade starts cold: run the eval set first, then cut over during low traffic. On Claude, changing the top-level effort between requests also invalidates the cache; use per-message effort (beta, on supported models) or pick one level per session. On OpenAI GPT-6 models, append a configuration_update input item instead of changing reasoning.effort.

Measure hit rate per endpoint

Read cache_read_input_tokens and cache_creation_input_tokens from Claude’s usage object (cached_tokens on OpenAI), and track the hit rate per endpoint, prompt version, and model. A sustained drop is a regression, usually from a new field at the top of the system message, unstable chunk ordering, a model change, or middleware inserting a request ID into the prompt. Alert on it; it surfaces faster than the cost dashboard.