Claude Haiku 5.5Latest
For high-volume, latency-sensitive tasks such as classification, extraction, and routing
- Context window
- 1Mtokens
- Max output
- 128Ktokens
- Input pricing
- From $0.10/ MTok
- Output pricing
- From $0.50/ MTok
Overview
Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.
For code changes, see the migration guide. For model IDs, pricing, and limits, see the Claude Haiku 5.5 overview. For prompting guidance, see Prompting Claude Haiku 5.5.
How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |
| Claude Opus 5.5 | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |
| Claude Sonnet 5.5 | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jun 2026 |
| Claude Haiku 5.5This model | 1M | 128K | From $0.10 / $0.50 | Fastest | Adaptive | medium | Jun 2026 |
Specifications
Model IDs
Pricing
- Input
- $0.10 / MTok for prompts up to 100,000 tokens$0.50 / MTok for prompts over 100,000 tokens
- Output
- $0.50 / MTok for prompts up to 100,000 tokens$2.50 / MTok for prompts over 100,000 tokens
- 5m cache write
- $0.125 / MTok for prompts up to 100,000 tokens$0.625 / MTok for prompts over 100,000 tokens
- 1h cache write
- $0.20 / MTok for prompts up to 100,000 tokens$1 / MTok for prompts over 100,000 tokens
- Cache read
- $0.01 / MTok for prompts up to 100,000 tokens$0.05 / MTok for prompts over 100,000 tokens
- Batch API
- 50% discount on input and output
Capabilities
- Context window
- 1M tokens
- Max output
- 128K tokens
- Max output (Batch API, beta)
- 300K tokens
- Thinking
- Adaptive
- Default effort
medium- Comparative latency
- Fastest
- Input → output
- Text and images → text
- Reliable knowledge cutoff
- Jun 2026
- Training data cutoff
- Jun 2026
Availability
- Status
- Active (latest)
- Released
- October 7, 2026
- Retirement
- Not sooner than October 7, 2027
- Platforms
- Claude APIAmazon BedrockGoogle CloudMicrosoft FoundryClaude Platform on AWS
Good to know
- Adaptive thinking is on by default. Control thinking depth with the effort parameter.
- Omit
temperature,top_p, andtop_k, since a non-default value for any of them returns a 400 error. - On the Message Batches API, Claude Haiku 5.5 supports up to 300k output tokens with the
output-300k-2026-03-24beta header. - Query limits and capabilities programmatically with the Models API.
Resources
Behavioral differences and prompting patterns specific to Claude Haiku 5.5.
Choose a model and effort level, shape prompts, and stream output for faster responses.
Claude Haiku 5.5 decides when and how much to think. Steer depth with effort.
1M tokens. How the window is counted and managed.
Reference
The system prompt Claude Haiku 5.5 uses on claude.ai and the Claude apps.
Safety evaluations and deployment decisions for Claude Haiku 5.5.
Full price list, including batch discounts and prompt caching rates.
How model IDs, aliases, and pinned snapshots work.
Lifecycle status and retirement commitments for every Claude model.
Was this page helpful?