Claude Sonnet 5.5Latest
The best combination of speed and intelligence
- Context window
- 1Mtokens
- Max output
- 128Ktokens
- Input pricing
- $2/ MTok
- Output pricing
- $10/ MTok
Overview
Claude Sonnet 5.5 offers the best combination of speed and intelligence. Five breaking changes affect code already running on Claude Sonnet 5:
- Turn off up-front thinking with
between_tools. - Forced tool use returns an error.
- Thinking blocks are tied to the model and the conversation.
- On the Claude API and Google Cloud, the earlier
computer_20251124computer use tool is not accepted. - The advisor tool rejects Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisors.
One more change alters the response shape without failing any request: text between tool calls comes back in thinking blocks. An application that streams that text to its users goes quiet between tool calls until it sets a display value that returns the text, or turns off up-front thinking with between_tools.
How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |
| Claude Opus 5.5 | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |
| Claude Sonnet 5.5This model | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jun 2026 |
| Claude Haiku 4.5 | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
Specifications
Model IDs
Pricing
- Input
- $2 / MTok
- Output
- $10 / MTok
- 5m cache write
- $2.50 / MTok
- 1h cache write
- $4 / MTok
- Cache read
- $0.20 / MTok
- Batch API
- 50% discount on input and output
Capabilities
- Context window
- 1M tokens
- Max output
- 128K tokens
- Max output (Batch API, beta)
- 300K tokens
- Thinking
- Adaptive
- Default effort
high- Comparative latency
- Fast
- Input → output
- Text and images → text
- Reliable knowledge cutoff
- Jun 2026
- Training data cutoff
- Jun 2026
Availability
- Status
- Active (latest)
- Released
- September 28, 2026
- Retirement
- Not sooner than September 28, 2027
- Platforms
- Claude APIAmazon BedrockGoogle CloudMicrosoft FoundryClaude Platform on AWS
Good to know
- Adaptive thinking is on by default. The lowest thinking setting is
between_tools, which turns off up-front thinking. It works athigheffort or below. See What's new in Claude Sonnet 5.5. - Setting
temperature,top_p, ortop_kto a non-default value returns a 400 error. - The minimum cacheable prompt length is 512 tokens. See Prompt caching.
- On the Message Batches API, Claude Sonnet 5.5 supports up to 300k output tokens with the
output-300k-2026-03-24beta header. - Query limits and capabilities programmatically with the Models API.
Resources
Behavioral differences and prompting patterns specific to Claude Sonnet 5.5.
The control for thinking depth, latency, and cost. Choose a level per workload.
How adaptive thinking works, which thinking settings each model accepts, and how thinking blocks are preserved.
Reference
Full price list, including batch discounts and prompt caching rates.
How model IDs, aliases, and pinned snapshots work.
Lifecycle status and retirement commitments for every Claude model.
Was this page helpful?