Claude Fable 5.1 keeps $10/$50 pricing but cuts cache reads 75% to $0.25/M tokens. Examples show 11–47% lower agent bills, plus benchmarks and safeguards.
Claude Fable 5.1 costs exactly what Fable 5 cost — $10 per million input tokens, $50 per million output — with one change that dominates every agent bill: cache reads dropped from $1.00 to $0.25 per million tokens. That is 2.5% of the input price, against the 10% multiplier every other Claude model uses. Anthropic’s own four weeks of August usage put the effect at roughly 25% lower bills for typical workloads and up to 45% for highly agentic ones. Released 1 September 2026 as claude-fable-5-1 on the Claude API, AWS, Google Cloud and Azure; Mythos 5.1 is the same model without production safeguards, restricted to vetted cybersecurity and life-sciences organisations.
The rest of this post is the arithmetic, because the percentage Anthropic quotes depends entirely on how much of your traffic is cache hits, and most teams have never measured that.
The price table that matters
| Model | Input | 5-min cache write | 1-hour cache write | Cache read | Output |
|---|---|---|---|---|---|
| Fable 5.1 / Mythos 5.1 | $10 | $12.50 | $20 | $0.25 | $50 |
| Fable 5 / Mythos 5 | $10 | $12.50 | $20 | $1.00 | $50 |
| Opus 5 / 4.8 | $5 | $6.25 | $10 | $0.50 | $25 |
| Sonnet 5 | $2 | $2.50 | $4 | $0.20 | $10 |
| Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | $5 |
All figures per million tokens from Anthropic’s pricing page as of 5 September 2026. Batch is 50% off input and output on every row. A second quiet change on the same page: Sonnet 5’s $2/$10 “introductory” price is now permanent — the increase to $3/$15 scheduled for 1 September did not happen.
Notice the strange result in the cache column: a Fable 5.1 cache read at $0.25 is now cheaper than a Sonnet 5 cache read at $0.20 in relative terms and nearly the same in absolute terms. For the re-read portion of an agent loop, the most expensive model and a mid-tier model cost about the same.
Three workloads, worked
Assume a coding agent with a 150,000-token stable prefix (repo map, CLAUDE.md, tool schemas), 40 turns per task, 3,000 fresh input tokens and 1,500 output tokens per turn. Prefix written once with a 5-minute cache, read on the other 39 turns.
| Cost component | Fable 5 | Fable 5.1 |
|---|---|---|
| Cache write, 150k × 1 turn | $1.875 | $1.875 |
| Cache reads, 150k × 39 turns = 5.85M | $5.85 | $1.46 |
| Fresh input, 3k × 40 = 120k | $1.20 | $1.20 |
| Output, 1.5k × 40 = 60k | $3.00 | $3.00 |
| Per task | $11.93 | $7.54 |
That is a 37% cut on a task shape that is ordinary for Claude Code and its imitators. Push the prefix to 400k tokens — a monorepo — and the same 40 turns go from $24.80 to $13.10, a 47% cut, because reads dominate everything. Pull the prefix down to 20k tokens (a chat product with a modest system prompt) and the saving shrinks to about 11%. Anthropic’s “25% typical, 45% agentic” range is consistent with those three points.
Run your own shape through the AI prompt cost calculator — it has cache-read and cache-write fields — before you decide anything. And if you want the actual ratio for your Claude Code sessions instead of an assumption, the new AI Chat Wrapped tool reads your local .jsonl logs and totals input, output, cache-write and cache-read tokens per model.
Comments · 0
Beta: comments are stored locally on your device and not visible to other readers.
No comments yet. Be the first to share your thoughts.