MartechARTICLE

Claude API pricing: why cache and tokenizer matter more than the chosen model for margin

The official Claude API pricing documentation, covering Haiku 4.5, Sonnet 5.5, Opus 5.5, and Fable 5.1, shows that cache, tokenizer, and fast mode weigh on margin as much as the choice of model.

What the pricing table says, in numbers

The official Claude API pricing documentation currently lists four production model families with a clearly defined price hierarchy: Claude Haiku 4.5, Claude Sonnet 5.5, Claude Opus 5.5, and Claude Fable 5.1. The gap between the cheapest and the most expensive tier is 10x, both for input and output tokens.

Official pricing table from Claude's documentation showing input, output, and cache values for the Fable 5.1 and Opus 5.5 models
Official pricing table from Claude's documentation showing input, output, and cache values for the Fable 5.1 and Opus 5.5 models. Reproduction: docs.anthropic.com.
ModelInputOutputCache hitUsage profile
Claude Haiku 4.5$1/MTok$5/MTok$0.10/MTokHigh volume, simple tasks
Claude Sonnet 5.5$2/MTok$10/MTok$0.20/MTokGeneral production, cost/quality balance
Claude Opus 5.5$4/MTok$20/MTok$0.20/MTokAgentic coding, long-running tasks
Claude Fable 5.1$10/MTok$50/MTok$0.25/MTokFrontier reasoning, long horizon

At first glance, this table is an obvious invitation to multi-model routing: why pay $10 per million input tokens on Fable 5.1 if Haiku 4.5 handles the same task for $1? But the full calculation, including prompt caching, Batch API, and fast mode, shows that the nominal list price is only the surface of the margin decision.

Cache narrows the gap between the expensive and the cheap model

Anthropic charges differently for cache writes and reads: a 5-minute write costs 1.25x the base input price, a 1-hour write costs 2x, and a hit (read) costs a fraction of the standard input price. Nothing surprising so far. The detail lies in the hit multiplier: 0.1x for most models, but 0.05x for Opus 5.5 and 0.025x for Fable 5.1.

Working out the absolute dollar cost per million tokens on a cache hit, the result is nearly flat across the four models: $0.10 on Haiku 4.5, $0.20 on Sonnet 5.5, $0.20 on Opus 5.5, and $0.25 on Fable 5.1. In other words, the 10x spread in nominal input price drops to 2.5x once the content is already cached.

In short: for a product with mostly static context (a long system prompt, a reused document base, conversation history with automatic caching), running on Fable 5.1 costs only slightly more than running on Haiku 4.5, because most tokens are read from cache rather than processed from scratch. This changes the logic of multi-model routing: the question is no longer

Translated from the Brazilian Portuguese original · Read the original