Claude Platform Docs
Models & pricingModels

Claude Sonnet 5.5Latest

The best combination of speed and intelligence

Try in playground
Context window
1Mtokens
Max output
128Ktokens
Input pricing
$2/ MTok
Output pricing
$10/ MTok

Overview

Claude Sonnet 5.5 offers the best combination of speed and intelligence. Five breaking changes affect code already running on Claude Sonnet 5:

One more change alters the response shape without failing any request: text between tool calls comes back in thinking blocks. An application that streams that text to its users goes quiet between tool calls until it sets a display value that returns the text, or turns off up-front thinking with between_tools.

What's new in Claude Sonnet 5.5

How it compares

ModelContextMax outputPrice / MTokLatencyThinkingDefault effortKnowledge cutoff
Claude Fable 5.11M128K$10 / $50SlowerAdaptive (always on)highJun 2026
Claude Opus 5.51M128K$4 / $20ModerateAdaptive (always on)mediumJun 2026
Claude Sonnet 5.5This model1M128K$2 / $10FastAdaptivehighJun 2026
Claude Haiku 4.5200K64K$1 / $5FastestExtended—Feb 2025

Specifications

Model IDs

Claude API
Amazon Bedrock
Google Cloud
Microsoft Foundry
Claude Platform on AWS

Pricing

Input
$2 / MTok
Output
$10 / MTok
5m cache write
$2.50 / MTok
Cache read
$0.20 / MTok
Batch API
50% discount on input and output

Full price list

Capabilities

Max output
128K tokens
Thinking
Adaptive
Comparative latency
Fast
Input → output
Text and images → text
Reliable knowledge cutoff
Jun 2026
Training data cutoff
Jun 2026

Availability

Status
Active (latest)
Released
September 28, 2026
Retirement
Not sooner than September 28, 2027

Good to know

  • Adaptive thinking is on by default. The lowest thinking setting is between_tools, which turns off up-front thinking. It works at high effort or below. See What's new in Claude Sonnet 5.5.
  • Setting temperature, top_p, or top_k to a non-default value returns a 400 error.
  • The minimum cacheable prompt length is 512 tokens. See Prompt caching.
  • On the Message Batches API, Claude Sonnet 5.5 supports up to 300k output tokens with the output-300k-2026-03-24 beta header.
  • Query limits and capabilities programmatically with the Models API.

Resources

Behavioral differences and prompting patterns specific to Claude Sonnet 5.5.

The control for thinking depth, latency, and cost. Choose a level per workload.

How adaptive thinking works, which thinking settings each model accepts, and how thinking blocks are preserved.

Reference

Full price list, including batch discounts and prompt caching rates.

How model IDs, aliases, and pinned snapshots work.

Lifecycle status and retirement commitments for every Claude model.

Was this page helpful?