3AM Marketer
AI

Claude Tokens: API Pricing, Caching, and Costs

Claude token costs vary tenfold across models, while prompt caching cuts input rates by 90 percent. Here is the operational math behind Anthropic token billing.

Claude Tokens: API Pricing, Caching, and Costs

Claude token pricing ranges from $1.00 to $10.00 per million input tokens across Anthropic's current model tier, while output tokens cost between $5.00 and $50.00 per million. Reusing cached prompt prefixes cuts input billing to one-tenth of standard rates, making cache hit efficiency the primary way to control costs in production pipelines.123

How Anthropic Measures Claude Tokens Across Model Tiers

Anthropic provides a dedicated token counting endpoint, client.messages.count_tokens, that calculates input tokens for messages, system prompts, client tools, and documents before sending calls to Claude. The endpoint accepts text, tool definitions, and base64-encoded images and PDFs.4

Token counts differ by tokenizer generation. Models starting from Claude 4.7 and Claude Mythos Preview use an updated tokenizer architecture. When converting the same raw input text, this newer tokenizer yields approximately 30 percent more tokens than earlier model releases. Anthropic explicitly recommends recounting prompts against the model you plan to use rather than reusing counts measured against earlier models.4

The Messages API usage object accounts for tokens across distinct buckets. You are billed strictly for your submitted content and generated responses. While Anthropic adds background tokens for certain system optimizations, documentation confirms that users are not billed for these system-added tokens.4

Claude Model Pricing Matrix: From Haiku 4.5 to Fable 5.1

Anthropic bills API usage per million tokens (MTok). Across current models, rates scale directly with model capacity and task complexity. Claude Haiku 4.5 provides the lowest latency and cost, Sonnet 5 handles balanced production workloads, Opus 5.5 processes long-running agentic coding, and Fable 5.1 targets demanding reasoning tasks.1

Anthropic API Token Pricing and Thinking Capabilities1
ModelAPI AliasInput Price (per MTok)Output Price (per MTok)Thinking Mode
Claude Fable 5.1claude-fable-5-1$10.00$50.00Adaptive (always on)
Claude Opus 5.5claude-opus-5-5$4.00$20.00Adaptive (always on)
Claude Sonnet 5claude-sonnet-5$2.00$10.00Adaptive
Claude Haiku 4.5claude-haiku-4-5-20251001$1.00$5.00Extended

Output generation carries a fivefold premium over standard input across all four current models. Generating one million tokens on Claude Haiku 4.5 costs $5.00, whereas generating one million tokens on Claude Fable 5.1 costs $50.00. Context windows span up to 1M tokens on Sonnet 5, Opus 5.5, and Fable 5.1, while Haiku 4.5 provides a 200K token context window.1

The Prompt Caching Discount: Slashing Input Costs by 90 Percent

Prompt caching optimizes API spend by allowing Claude to resume processing from static prefixes in your prompts. Rather than reprocessing an entire context on every request, the platform stores processed prefixes. On most models, cache read operations are billed at 10 percent of the standard input token rate, securing a 90 percent discount for cached tokens.23

On Claude Sonnet 5, standard input tokens cost $2.00 per million. When read from cache, those input tokens cost $0.20 per million. On Claude Opus 5.5, input costs drop from $4.00 per million to $0.20 per million. In high-turn agent sessions and document retrieval workflows, cached tokens often represent the overwhelming majority of network payload. Log analysis of heavy Claude Code workflows indicates that cache reads can comprise up to 96.9 percent of total token volume.123

Cache entries have a default lifetime of five minutes. Anthropic calculates this duration from the beginning of the request that creates or reads the cache entry, not when the response completes. If an agent takes four minutes to finish generation, a follow-up request that reuses the same cached prefix must start within approximately one minute of that response completing. The cache refreshes for no additional cost each time the cached content is used.2

Worked Cost Example: 20-Turn Agent Session on Sonnet 5

Consider an interactive agent running on Claude Sonnet 5 with a static 20,000-token system instruction block containing tool definitions and guidelines. Over a 20-turn session, the agent processes that 20,000-token prefix on every single turn.12

Without prompt caching, the agent re-evaluates the 20,000 tokens on all 20 turns, totaling 400,000 standard input tokens. At $2.00 per million tokens, processing the static instruction context costs $0.80. With prompt caching enabled, turn one creates the cache entry at standard base rates. Turns 2 through 20 read the prefix from cache at $0.20 per million tokens. The initial turn costs $0.04 (20,000 tokens at $2.00/MTok), while the subsequent 19 turns cost $0.076 combined (380,000 tokens at $0.20/MTok). The total prefix cost drops from $0.80 down to $0.116, saving over 85 percent on context ingestion.12

Extended and Adaptive Thinking: Managing Reasoning Tokens

Anthropic models use internal reasoning capabilities that generate thinking tokens before producing user-visible text. In legacy extended thinking mode, developers configure a manual thinking target using the budget_tokens parameter inside the API request. Claude consumes internal tokens against that designated budget to reason through complex logic before streaming its final response.6

Manual extended thinking using budget_tokens is deprecated on Claude 4.6 models and completely rejected on Claude 4.7 and later releases with a 400 error. Modern tiers like Claude Sonnet 5 and Claude Opus 5.5 utilize adaptive thinking modes, while Claude Fable 5.1 maintains adaptive thinking as an always-on feature. Claude Haiku 4.5 retains support for extended thinking.16

Set strict max_tokens limits to prevent runaway reasoning loops from inflating your bill. Because reasoning tokens are metered alongside generated content, unconstrained thinking blocks consume output context allowances and increase total request fees.16

Why Agent Loops and Claude Code Experience Invisible Token Burn

Developers running automated terminal tools like Claude Code frequently report unexpected billing spikes. These spikes are rarely caused by user prompt length. Instead, they stem from agentic context assembly rules that inadvertently break prompt caching.35

Prompt caching operates strictly from the beginning of the request forward. The system caches content sequentially up to designated breakpoints. If any dynamic content changes early in the payload, the entire downstream prefix misses the cache. Developer issue reports show that injecting dynamic git status strings and changing recent commit messages into early instructions causes skills and project configuration files to miss cache breakpoints on every turn.25

a git status + recent commits (that will always change) and a missing cache-mark that will make skills & project-claude.md cachemiss every time too

Hacker News Technical Discussion on Claude Code Token Usage

When cache misses occur continuously, multi-turn agent sessions reprocess the entire accumulated conversational history at full base input prices instead of discounted cache read prices. Disabling fluctuating dynamic headers or restructuring where runtime state is appended restores cache reuse rates above 95 percent.35

Production Decision Framework: Sub-Agents and Caching Audits

To stop bill inflation, apply the cost optimization methods outlined in Anthropic's developer resources, which focus on finding the Pareto-optimal balance between task pass rate and cost per task:7

Start by auditing cache hits using API response headers. If cache read tokens account for less than 80 percent of your multi-turn input volume, prompt layout is invalidating the cache prefix. Move all dynamic environmental data, user variables, and runtime git logs below your static system prompts and tool definitions.235

Next, implement sub-agent architectures that combine Haiku with Opus. Route high-volume retrieval, initial triage, and lightweight formatting to Claude Haiku 4.5 at $1.00 per million input tokens. Reserve Claude Opus 5.5 ($4.00/MTok) and Claude Fable 5.1 ($10.00/MTok) exclusively for complex multi-step reasoning. Finally, review whether your team requires Pro, Team, or Enterprise subscription plans or direct API billing to match your organization's workload patterns.178

Questions

How much is 1 token in Claude?

A token is a fragment of text processed by the model. On Claude Haiku 4.5, one standard input token costs $0.000001, while on Claude Sonnet 5, one input token costs $0.000002. Reusing cached input tokens drops these costs by 90 percent.

How much is 1 million tokens in Claude?

One million input tokens costs $1.00 on Claude Haiku 4.5, $2.00 on Claude Sonnet 5, $4.00 on Claude Opus 5.5, and $10.00 on Claude Fable 5.1. Output tokens cost five times the input rate across all four models, ranging from $5.00 to $50.00 per million.

What are tokens in Claude?

Tokens are the numerical units that Claude uses to process text, code, and structured data. Anthropic's newer tokenizer for Claude 4.7 and later models produces approximately 30 percent more tokens on identical text compared to earlier model generations.

How many tokens does $5 get you in Claude?

A $5.00 balance purchases 5 million standard input tokens on Claude Haiku 4.5 or 2.5 million input tokens on Claude Sonnet 5. When reading fully from prompt cache prefixes, $5.00 pays for 25 million cached input tokens on Sonnet 5.

What we got wrong

Corrected the cached input price for Claude Opus 5.5 from $0.40 to $0.20 per million tokens, and noted that the 10 percent cache read rate applies to most models, not all.

Anthropic prices cache hits on Claude Opus 5.5 at 0.05x the base input price, not the standard 0.1x, and Claude Fable 5.1 at 0.025x.