Claude Platform Docs
Messages模型功能

快速模式(研究預覽版)

從支援的 Claude Opus 模型獲得最高 2.5 倍的每秒輸出 token 數。

「Fast mode」(快速模式)以進階定價,為 Claude Opus 5.5、Claude Opus 5 和 Claude Opus 4.8 提供最高 2.5 倍的「output tokens per second」(每秒輸出 token 數)。若要啟用,請在您的請求中設定 speed: "fast" 並搭配 fast-mode-2026-02-01 beta 標頭。

支援的模型

快速模式支援以下模型:

  • Claude Opus 5.5()
  • Claude Opus 5()
  • Claude Opus 4.8()

快速模式的運作方式

快速模式以更快的推論配置執行相同的模型。智慧或能力沒有任何改變。

  • 與標準速度相比,每秒輸出 token 數最高提升 2.5 倍
  • 速度優勢集中於「output tokens per second」(每秒輸出 token 數),即 OTPS,而非「time to first token」(首個 token 時間),即 TTFT
  • 相同的模型權重與行為(並非不同的模型)
  • 與 streaming(串流)相容,在串流中 OTPS 的提升最為明顯

基本用法

client = anthropic.Anthropic()

response = client.beta.messages.create(
    model="claude-opus-5-5",
    max_tokens=4096,
    speed="fast",
    betas=["fast-mode-2026-02-01"],
    messages=[
        {"role": "user", "content": "Refactor this module to use dependency injection"}
    ],
)

for block in response.content:
    if block.type == "text":
        print(block.text)

定價

快速模式的定價是在整個 context window(上下文視窗)範圍內以標準費率的倍數計算,包括超過 200k 輸入 token 的請求。下表顯示支援模型的快速模式定價:

模型輸入輸出
Claude Opus 5.5$8 USD / MTok$40 USD / MTok
Claude Opus 5 / Claude Opus 4.8$10 USD / MTok$50 USD / MTok

快速模式定價可與其他定價調整因子疊加:

如需完整的定價詳情,請參閱定價頁面。

速率限制

快速模式擁有專屬的 rate limit(速率限制),與標準 Opus 速率限制分開。當您的快速模式速率限制被超過時,API 會回傳 429 錯誤,並附帶 retry-after 標頭,指示何時會有可用容量。

回應中包含指示您快速模式速率限制狀態的標頭:

標頭說明
anthropic-fast-input-tokens-limit每分鐘快速模式輸入 token 的上限
anthropic-fast-input-tokens-remaining剩餘的快速模式輸入 token
anthropic-fast-input-tokens-reset快速模式輸入 token 限制重設的時間
anthropic-fast-output-tokens-limit每分鐘快速模式輸出 token 的上限
anthropic-fast-output-tokens-remaining剩餘的快速模式輸出 token
anthropic-fast-output-tokens-reset快速模式輸出 token 限制重設的時間

如需各層級的速率限制,請參閱速率限制頁面。

檢查使用了哪種速度

回應的 usage 物件包含一個 speed 欄位,指示使用了哪種速度,值為 "fast" 或 "standard"。在不支援快速模式的模型上請求 speed: "fast" 會回傳錯誤,超過快速模式的速率限制或容量(429 或 529)也同樣會回傳錯誤。當帶有 speed: "fast" 的請求成功時,usage.speed 為 "fast"。如果您使用 Claude Opus 4.6 並請求快速模式,其行為是獨特的。它不會像其他不支援快速模式的模型那樣回傳錯誤,而是靜默地切換至標準速度。雖然 Opus 4.6 不會產生錯誤,但 speed 欄位會準確地顯示 "standard"。

client = anthropic.Anthropic()

response = client.beta.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    speed="fast",
    betas=["fast-mode-2026-02-01"],
    messages=[{"role": "user", "content": "Hello"}],
)

print(response.usage.speed)  # "fast" or "standard"
Output
{
  "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
  "type": "message",
  "role": "assistant",

  "usage": {
    "input_tokens": 8,
    "output_tokens": 12,
    "speed": "fast"
  }
}

若要追蹤整個組織的快速模式使用量與成本,請參閱使用量與成本 API。

重試與退回

自動重試

當超過快速模式速率限制時,API 會傳回 429 錯誤,並附上 retry-after 標頭。Anthropic SDK 預設會自動重試這些請求最多 2 次(可透過 max_retries 設定),並在每次重試前等待伺服器指定的延遲時間。由於快速模式採用持續的 token 補充機制,retry-after 延遲通常很短,一旦有可用容量,請求便會成功。

退回至標準速度

如果您希望退回標準速度,而不是等待快速模式的容量,請捕捉速率限制錯誤,並在不帶 speed: "fast" 的情況下重試。在初始的快速請求上將 max_retries 設為 0,即可略過自動重試,並在發生速率限制錯誤時立即失敗。

由於將 max_retries 設為 0 也會停用其他暫時性錯誤(過載、內部伺服器錯誤)的重試,以下範例會針對這些情況,以預設重試設定重新發出原始請求。

client = anthropic.Anthropic()


def create_message_with_fast_fallback(max_retries=0, max_attempts=3, **params):
    try:
        return client.with_options(max_retries=max_retries).beta.messages.create(
            **params
        )
    except anthropic.RateLimitError:
        if params.get("speed") == "fast":
            del params["speed"]
            return create_message_with_fast_fallback(max_retries=max_retries, **params)
        raise
    except (
        anthropic.APIStatusError,
        anthropic.APIConnectionError,
    ) as error:
        if isinstance(error, anthropic.APIStatusError) and error.status_code < 500:
            raise
        if max_attempts > 1:
            return create_message_with_fast_fallback(
                max_retries=max_retries, max_attempts=max_attempts - 1, **params
            )
        raise


message = create_message_with_fast_fallback(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
    betas=["fast-mode-2026-02-01"],
    speed="fast",
    max_retries=0,
)

注意事項

  • 提示快取: 在快速與標準速度之間切換會使提示快取失效。不同速度的請求不會共用快取的前綴。
  • 支援的模型: 快速模式支援 Claude Opus 5.5、Claude Opus 5 和 Claude Opus 4.8。請參閱支援的模型。
  • TTFT: 快速模式的優勢集中在每秒輸出 token 數(OTPS),而非首個 token 產生時間(TTFT)。
  • Batch API: 快速模式不適用於 Batch API。
  • Priority Tier: 快速模式不適用於 Priority Tier 承諾方案。
  • Claude Platform on AWS: 快速模式目前不適用於 Claude Platform on AWS。

後續步驟

從代理工作流程中取得經過驗證的 JSON 結果。

了解 Anthropic 針對模型與功能的定價結構。

使用 effort 參數控制 Claude 回應時使用的 token 數量,在回應的完整性與 token 效率之間取得平衡。

透過伺服器傳送事件(server-sent events)以增量方式串流 Messages API 回應,包括文字、工具使用與擴展思考的增量內容。

Was this page helpful?