令牌计数让您可以在将消息发送给 Claude 之前确定消息中的令牌数量。这有助于您对提示和使用情况做出明智的决策。通过令牌计数,您可以:
此功能符合零数据保留(ZDR)的条件。当您的组织签订了 ZDR 协议时,通过此功能发送的数据在 API 响应返回后不会被存储。
令牌计数端点接受与创建消息相同的结构化输入列表,包括对系统提示、工具、图像和 PDF 的支持。响应包含输入令牌的总数。
令牌计数应被视为一个估计值。在某些情况下,创建消息时实际使用的输入令牌数量可能会有少量差异。
令牌计数可能包括 Anthropic 为系统优化而自动添加的令牌。系统添加的令牌不会向您收费。计费仅反映您的内容。
所有活跃模型都支持令牌计数,包括 Claude Sonnet 5。
Claude Opus 4.7 及更高版本的 Opus 模型、Claude Fable 5、Claude Mythos 5、Claude Mythos Preview 和 Claude Sonnet 5 使用较新的分词器。相同的输入文本产生的令牌数量比早期模型多约 30%。确切的增幅取决于内容和工作负载的形态。请针对您计划使用的模型重新计算提示的令牌数,而不是重复使用针对早期模型测量的计数。
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-opus-4-8",
system="You are a scientist",
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(response.json()){ "input_tokens": 14 }服务器工具的令牌计数仅适用于第一次采样调用。
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-opus-4-8",
tools=[
{
"name": "get_weather",
"description": "Get the current weather in a given location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA",
}
},
"required": ["location"],
},
}
],
messages=[{"role": "user", "content": "What's the weather like in San Francisco?"}],
)
print(response.json()){ "input_tokens": 403 }import base64
import httpx
image_url = "https://upload.wikimedia.org/wikipedia/commons/a/a7/Camponotus_flavomarginatus_ant.jpg"
image_media_type = "image/jpeg"
image_data = base64.standard_b64encode(httpx.get(image_url).content).decode("utf-8")
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-opus-4-8",
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": image_media_type,
"data": image_data,
},
},
{"type": "text", "text": "Describe this image"},
],
}
],
)
print(response.json()){ "input_tokens": 1551 }请参阅使用扩展思考时如何计算上下文窗口了解更多详情
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-sonnet-4-6",
thinking={"type": "enabled", "budget_tokens": 16000},
messages=[
{
"role": "user",
"content": "Are there an infinite number of prime numbers such that n mod 4 == 3?",
},
{
"role": "assistant",
"content": [
{
"type": "thinking",
"thinking": "This is a nice number theory question. Let's think about it step by step...",
"signature": "EuYBCkQYAiJAgCs1le6/Pol5Z4/JMomVOouGrWdhYNsH3ukzUECbB6iWrSQtsQuRHJID6lWV...",
},
{
"type": "text",
"text": "Yes, there are infinitely many prime numbers p such that p mod 4 = 3...",
},
],
},
{"role": "user", "content": "Can you write a formal proof?"},
],
)
print(response.json()){ "input_tokens": 88 }令牌计数支持 PDF,其限制与 Messages API 相同。
import base64
import anthropic
client = anthropic.Anthropic()
with open("/path/to/document.pdf", "rb") as pdf_file:
pdf_base64 = base64.standard_b64encode(pdf_file.read()).decode("utf-8")
response = client.messages.count_tokens(
model="claude-opus-4-8",
messages=[
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": pdf_base64,
},
},
{"type": "text", "text": "Please summarize this document."},
],
}
],
)
print(response.json()){ "input_tokens": 2188 }Claude Fable 5 和 Claude Mythos 5 使用随 Claude Opus 4.7 引入的分词器,对于相同的文本,其产生的令牌数量比 Claude Opus 4.7 之前的模型多约 30%。确切的增幅取决于内容和工作负载的形态。令牌计数端点会根据您传入的 model 所使用的分词器返回计数,因此要测量您的工作负载的差异,请对同一请求计数两次:一次使用您当前的模型,一次使用 model: "claude-fable-5"(或 "claude-mythos-5"),然后比较两个 input_tokens 值。
计费和迁移: Claude Fable 5 和 Claude Mythos 5 上的使用量和计费反映了此分词器的计数。如果您从 Claude Opus 4.7 之前的模型迁移,相同的内容会消耗大约多 30% 的令牌。确切的增幅取决于内容和工作负载的形态。将工作负载迁移到 Claude Fable 5 和 Claude Mythos 5 时,请勿重复使用在 Claude Opus 4.7 之前的模型上测量的令牌计数来估算成本或上下文窗口的适配情况。请使用 model: "claude-fable-5"(或 "claude-mythos-5")来计算您的提示的令牌数。
令牌计数免费使用,但受基于您的使用层级的每分钟请求数速率限制约束。如果您需要更高的限制,请在限制页面上使用请求提高速率限制。
| 使用层级 | 每分钟请求数 (RPM) |
|---|---|
| Start | 2,000 |
| Build | 4,000 |
| Scale | 8,000 |
令牌计数和消息创建具有各自独立的速率限制。使用其中一个不会计入另一个的限制。
阅读令牌计数端点的完整 API 参考。
使用令牌计数将提示保持在模型的上下文窗口内。
在发送请求之前检查令牌计数,以保持在您的使用层级内。
通过缓存提示前缀来降低重复提示的成本和延迟。
Was this page helpful?