Haiku 5.5 costs $0.10 and $0.50 per million tokens under 100K, then $0.50 and $2.50 above. See the budget for 1M agent calls and how to stay in tier.
Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill. Anthropic released the model on 7 October 2026 on its own platform, AWS, Google Cloud and Microsoft Azure the same day. Output is $0.50 per million tokens at or below the threshold and $2.50 above it. The context window is 1M tokens with up to 128K tokens of output.
Short answer: for high-volume agent loops, Haiku 5.5 is the cheapest Anthropic model ever listed, and a loop that keeps every prompt under 100K tokens runs about $1,050 per million calls at 8,000 input tokens and 500 output tokens each. Add prompt caching and that falls to roughly $520. The risk is the 100K line. One careless context-stuffing change moves a call from the cheap tier to the expensive one. Here is the price sheet, the worked budget and the rules that keep you below the line.
The two-tier price sheet
The numbers below come from Anthropic's pricing page as reported by VentureBeat and MarkTechPost on launch day. Anthropic describes Haiku 5.5 as its "best lightweight model yet" and says it beats Haiku 4.5 in coding, computer use and knowledge work. It also says the model is on average 75% cheaper than Haiku 4.5. VentureBeat frames the headline as a 90% API price cut against Haiku 4.5's $1.00 input and $5.00 output, which is what you get when you compare the lower tier alone.
| Item | Prompt up to 100K tokens | Prompt above 100K tokens |
|---|---|---|
| Input, per million tokens | $0.10 | $0.50 |
| Output, per million tokens | $0.50 | $2.50 |
| Cache read, per million tokens | $0.01 | not listed in my sources |
| Cache write (5 minute), per million tokens | $0.125 | not listed in my sources |
Read the two columns as a step function, not a slope. A 99,000-token prompt with a 1,000-token answer costs $0.0104. A 101,000-token prompt with the same answer costs $0.053, which is 5.1 times more for 2% more input. That calculation assumes the higher rate applies to the whole request, which is how the tier is listed. Confirm on Anthropic's pricing page whether your region and provider bill it that way before you build a forecast on it.
The cache read price is worth a second look: $0.01 per million is one tenth of the base input price. If most of your prompt is a fixed system prompt, tool definitions and few-shot examples, you pay a tenth for those tokens on every hit. The cache write costs $0.125 per million, a 25% premium over base input, charged when the prefix is first stored.
A worked budget for one million calls
Assume a classification or routing agent loop. Each call sends 8,000 input tokens and returns 500 output tokens. Of the input, 6,000 tokens are a fixed prefix: instructions, tool schemas and examples. The other 2,000 vary per call. All numbers are mine, for illustration. Swap in your own.
| Scenario (1M calls, 8,000 in, 500 out) | Input cost | Output cost | Total |
|---|---|---|---|
| Haiku 5.5, no caching | 8,000M tokens x $0.10 = $800 | 500M x $0.50 = $250 | $1,050 |
| Haiku 5.5, 6,000-token prefix cached | 6,000M x $0.01 + 2,000M x $0.10 = $260 | $250 | $510, plus about $7.50 of cache writes |
| Haiku 4.5 ($1.00 in, $5.00 out) | $8,000 | $2,500 | $10,500 |
| Gemini 3.5 Flash ($1.50 in, $9.00 out) | $12,000 | $4,500 | $16,500 |
The cache write line assumes one prefix write per 100 calls, which is 10,000 writes of 6,000 tokens at $0.125 per million. That is a hot-cache assumption. If your traffic is bursty and the five-minute cache expires between bursts, you pay the write more often, and the saving shrinks.
The ratio is the point. Under identical traffic Haiku 5.5 comes out near one tenth of Haiku 4.5 and under one fifteenth of Gemini 3.5 Flash at its listed $1.50 and $9.00. Gemini 3.5 Flash launched on 19 May 2026 with a cached price of $0.15 and batch pricing of $0.75 and $4.50, so its gap narrows with batch jobs but does not close. A newer Gemini 3.8 Flash is reported at an introductory $0.75 and $3.75 through 31 December 2026, then $1.50 and $7.50. Sources disagree on those figures, so check Google's page before relying on them.
OpenAI's GPT-6 Luna is the direct competitor. VentureBeat describes Haiku 5.5 as priced in line with it. I do not have a verified Luna price sheet from this week, so I am not putting a Luna row in the table. Price your own workload on the AI Model Cost Calculator with both models side by side.
Comments · 0
Beta: comments are stored locally on your device and not visible to other readers.
No comments yet. Be the first to share your thoughts.