Claude Platform Docs

Claude Haiku 5.5Latest

For high-volume, latency-sensitive tasks such as classification, extraction, and routing

Try in playground
Context window
1Mtokens
Max output
128Ktokens
Input pricing
From $0.10/ MTok
Output pricing
From $0.50/ MTok

Overview

Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.

For code changes, see the migration guide. For model IDs, pricing, and limits, see the Claude Haiku 5.5 overview. For prompting guidance, see Prompting Claude Haiku 5.5.

What's new in Claude Haiku 5.5

How it compares

ModelContextMax outputPrice / MTokLatencyThinkingDefault effortKnowledge cutoff
Claude Fable 5.11M128K$10 / $50SlowerAdaptive (always on)highJun 2026
Claude Opus 5.51M128K$4 / $20ModerateAdaptive (always on)mediumJun 2026
Claude Sonnet 5.51M128K$2 / $10FastAdaptivehighJun 2026
Claude Haiku 5.5This model1M128KFrom $0.10 / $0.50FastestAdaptivemediumJun 2026

Specifications

Model IDs

Claude API
Amazon Bedrock
Google Cloud
Microsoft Foundry
Claude Platform on AWS

Pricing

Input
$0.10 / MTok for prompts up to 100,000 tokens$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens$2.50 / MTok for prompts over 100,000 tokens
5m cache write
$0.125 / MTok for prompts up to 100,000 tokens$0.625 / MTok for prompts over 100,000 tokens
1h cache write
$0.20 / MTok for prompts up to 100,000 tokens$1 / MTok for prompts over 100,000 tokens
Cache read
$0.01 / MTok for prompts up to 100,000 tokens$0.05 / MTok for prompts over 100,000 tokens
Batch API
50% discount on input and output

Full price list

Capabilities

Max output
128K tokens
Thinking
Adaptive
Comparative latency
Fastest
Input → output
Text and images → text
Reliable knowledge cutoff
Jun 2026
Training data cutoff
Jun 2026

Availability

Status
Active (latest)
Released
October 7, 2026
Retirement
Not sooner than October 7, 2027

Good to know

  • Adaptive thinking is on by default. Control thinking depth with the effort parameter.
  • Omit temperature, top_p, and top_k, since a non-default value for any of them returns a 400 error.
  • On the Message Batches API, Claude Haiku 5.5 supports up to 300k output tokens with the output-300k-2026-03-24 beta header.
  • Query limits and capabilities programmatically with the Models API.

Resources

Behavioral differences and prompting patterns specific to Claude Haiku 5.5.

Choose a model and effort level, shape prompts, and stream output for faster responses.

Claude Haiku 5.5 decides when and how much to think. Steer depth with effort.

1M tokens. How the window is counted and managed.

Reference

The system prompt Claude Haiku 5.5 uses on claude.ai and the Claude apps.

Safety evaluations and deployment decisions for Claude Haiku 5.5.

Full price list, including batch discounts and prompt caching rates.

How model IDs, aliases, and pinned snapshots work.

Lifecycle status and retirement commitments for every Claude model.

Was this page helpful?