Claude Prompt Caching: TTL Multipliers and Pricing
Claude prompt caching reduces repetitive input costs by up to 95 percent using 5-minute and 1-hour write tiers across supported models.
Claude prompt caching optimizes API workloads by allowing requests to resume from designated prompt prefixes instead of reprocessing full conversational contexts. Writing to the default 5-minute cache costs 1.25 times the base input token price, while reading from an established cache cuts costs to 0.1 times the base rate on most models and 0.05 times on Claude Opus 5.5 and Claude Sonnet 5.5.1
Understanding cache lifetime resets and token thresholds prevents costly architectural missteps. Developers evaluating API overhead can contextualize these mechanics alongside the broader Claude tokens pricing and cost guide.1
Prefixes, rolling TTL, and automatic breakpoints
Anthropic supports prompt caching across all active Claude models using two implementation modes: automatic caching and explicit breakpoints. Automatic caching requires adding a single cache_control field at the top level of the API request, which directs the system to assign the breakpoint to the last cacheable block and move it forward as turns advance.1

Explicit cache breakpoints require placing the cache_control block directly on individual content blocks to define precise boundary points. Prompt caching evaluates prompts in a fixed prefix sequence across designated blocks up to and including the block marked with cache_control:1
- tools
- system
- messages1
For workflows with longer intervals between requests, Anthropic provides an optional 1-hour cache duration at an additional cost. Writing to this extended tier costs 2 times the base input token price, compared to 1.25 times for the default 5-minute tier.1
Unit pricing across Claude models and cache tiers
Prompt cache write operations incur a surcharge over base input rates, but cache read hits generate steep discounts. Standard cache reads cost 0.1 times base input prices across most catalog tiers, but Claude Opus 5.5 and Claude Sonnet 5.5 reduce that factor to 0.05 times the base input price. These multipliers stack with other contractual adjustments, including Batch API discounts and data residency modifiers.1
For teams deploying multi-model pipelines, the spread in unit costs between entry-tier and flagship models remains pronounced. From our data: For 1 million input plus 1 million output tokens, Anthropic Claude Haiku 3.5 costs $4.8 and Anthropic Claude Opus 4 costs $90, 18.8 times as much.4
On Claude Sonnet 5, base input tokens cost $2 per million, 5-minute writes cost $2.50 per million, 1-hour writes cost $4 per million, and cache reads cost $0.20 per million. On Claude Opus 5, base input tokens cost $5 per million, 5-minute writes cost $6.25 per million, 1-hour writes cost $10 per million, and cache reads cost $0.50 per million. Comparing high-end deployments reveals distinct rates: a deeper dive is covered in the Claude Opus 5.5 vs Sonnet 5.5 analysis.13
On Claude Haiku 4.5, base input tokens cost $1 per million, 5-minute writes cost $1.25 per million, 1-hour writes cost $2 per million, and cache reads cost $0.10 per million. Older models reflect previous base structures: Claude Haiku 3.5 charges $0.80 per million base input, $1 per million for 5-minute writes, $1.60 per million for 1-hour writes, and $0.08 per million for reads. Claude Sonnet 4 charges $3 per million base input, $3.75 per million for 5-minute writes, $6 per million for 1-hour writes, and $0.30 per million for reads. Claude Opus 4 charges $15 per million base input, $18.75 per million for 5-minute writes, $30 per million for 1-hour writes, and $1.50 per million for reads.1
Minimum token thresholds and the silent failure rule
A prompt prefix must satisfy strict token length floors before Anthropic infrastructure will write it to the cache. On Amazon Bedrock, Claude Haiku 5.5, Claude Sonnet 5.5, and Claude Opus 5.5 require a minimum prefix length of 512 tokens per checkpoint. In contrast, Claude Haiku 4.5 on Amazon Bedrock requires a minimum prefix of 4,096 tokens per checkpoint.2
Direct API endpoints for Claude Sonnet 5 and Claude Opus 4.8 enforce a minimum prefix threshold of 1,024 tokens before caching activates. Requests that attach cache_control to blocks falling below these model thresholds do not fail or return an API error. The request executes successfully, but the prefix is not cached, causing every subsequent request to process at the full base input price.23
Prefix invalidation rules and 4-checkpoint request limits
Prompt caching architectures must operate within structural constraints to prevent accidental cache misses. Supported Claude models accept a maximum of 4 cache checkpoints or writes per individual request. On Amazon Bedrock, prompt cache checkpoints are accepted in system, messages, and tools fields.2
Prefix evaluation demands exact matching from the start of the prompt. Because evaluation processes tools before subsequent content, changing a single tool definition invalidates cached prefixes across tools, system prompts, and message blocks simultaneously. Toggling a narrower setting, such as web search, invalidates only the tools block prefix.13
When the caching engine walks backward through a conversation history to match an existing cache entry, it checks at most 20 block positions per breakpoint. Furthermore, when configuring requests that combine both 1-hour and 5-minute cache durations, any 1-hour cache block must appear before any 5-minute cache block in the sequence.3
Tracking context windows and token endpoints
Total context window calculations encompass all active tokens generated and consumed during execution. Context accounting tracks input and output tokens alongside thinking tokens, while tracking all cached volume across input_tokens, cache_read_input_tokens, and cache_creation_input_tokens.3
To prevent sub-threshold caching failures and context exhaustion, Anthropic maintains a dedicated token-counting endpoint. This endpoint allows engineers to measure prompt block sizes before dispatching requests, operating under its own separate rate limit from primary message generation endpoints.3
Deciding between cache configurations comes down to request frequency. Workflows executing calls more frequently than once every 5 minutes should maintain the default 5-minute TTL, using rolling zero-cost refreshes to avoid paying the 2.0x 1-hour write premium. If gaps between requests consistently exceed 5 minutes but remain within an hour, switching to the 1-hour TTL avoids paying full base input rates on subsequent turns, provided tool definitions remain completely static to prevent multi-block invalidation.13