Claude Platform Docs
Models & pricingClaude Haiku 5.5

What's new in Claude Haiku 5.5

Overview of new capabilities, breaking changes, and behavior changes in Claude Haiku 5.5, with a link to each feature's guide and to the migration guide for code changes.

Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5. Its thinking blocks work only in the account that produced them, or in an account linked to it.

For code changes, see the migration guide. For model IDs, pricing, and limits, see the Claude Haiku 5.5 overview. For prompting guidance, see Prompting Claude Haiku 5.5.

Summary of changes from Claude Haiku 4.5

Each row names one change, whether it is new, changed, or breaking, and what your code has to do.

ChangeTypeAction needed
Adaptive thinking and effortNewOptional: set effort to trade response quality against speed and cost.
Larger context window and outputNewNone. Existing max_tokens values stay valid, but thinking tokens count toward them.
Browser use toolNewNone. Available on the Claude API and Google Cloud.
Safety classifiers can decline a requestNewHandle stop_reason: "refusal" in your client. Server-side fallback isn't available.
Manual extended thinking returns an errorBreakingReplace budget_tokens with adaptive thinking.
Non-default sampling parameters return an errorBreakingOmit temperature, top_p, and top_k.
Assistant message prefill returns an errorBreakingEnd messages with a user turn.
Computer use needs the toolset on the Claude API and Google CloudBreakingReplace computer_20250124 with computer_toolset_20260801.
Changing earlier turns invalidates thinking blocksBreakingKeep conversations append-only if you send thinking blocks back.
Responses can begin with thinking blocksChangedSelect content blocks by type, not by position.
Thinking text is omitted by defaultChangedTo receive summarized thinking, set thinking.display to "summarized".
Same text counts as more tokensChangedRecount prompts and revisit max_tokens and cost estimates.
Replaying thinking blocks across accountsChangedIf you replay stored conversations through a different account, replay each through the account that produced it.

New capabilities

Adaptive thinking and effort

With adaptive thinking, Claude Haiku 5.5 decides when and how much to think. Adaptive thinking is on by default. While you can still turn thinking off with thinking: {"type": "disabled"} at high effort or below, the better way to trade response quality against speed and cost is to use the effort parameter.

Larger context window and output

Claude Haiku 5.5 has a 1M token context window and returns up to 128k output tokens, up from 200k and 64k on Claude Haiku 4.5. Existing max_tokens values stay valid, but thinking tokens count toward max_tokens, so a small limit can stop after a thinking block and before any text. See Configure thinking.

Behavior changes

Responses can begin with thinking blocks

Adaptive thinking is on by default, so a response can begin with one or more thinking blocks even when the request doesn't mention thinking. Code that reads the first content block as the answer needs to select blocks by their type field. See Configure thinking in the migration guide.

Same text counts as more tokens

Claude Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later models. As with all models that use this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5. The exact increase depends on the content. The shape of requests and responses doesn't depend on the tokenizer, but anything you measure or budget in tokens changes. See Recount tokens in the migration guide.

Replaying thinking blocks across accounts

Thinking blocks from Claude Haiku 5.5 work only in the account that produced them, or in an account linked to it. This matters only if you store conversations and replay them through a different account. See Replay thinking blocks through the account that produced them.

Was this page helpful?