编程式工具调用
让 Claude 在代码执行容器中通过代码调用您的工具,减少多工具工作流中的模型往返次数和令牌使用量。
"Programmatic tool calling"(编程式工具调用)允许 Claude 编写代码,在代码执行容器中以编程方式调用您的工具,而无需为每次工具调用都经过模型往返。这降低了多工具工作流的 "latency"(延迟),并通过允许 Claude 在数据进入模型的 "context window"(上下文窗口)之前对其进行过滤或处理来减少令牌消耗。在 BrowseComp 和 DeepSearchQA 等测试多步骤网络研究和复杂信息检索的智能体搜索基准测试中,在基础搜索工具之上添加编程式工具调用使性能平均提升了 11%,同时输入令牌使用量减少了 24%(请参阅通过动态过滤改进网络搜索)。
以检查 20 名员工的预算合规情况为例:传统方法需要 20 次独立的模型往返,并在此过程中将数千条费用明细拉入上下文。使用编程式工具调用,单个脚本即可运行全部 20 次查询、过滤结果,并仅返回超出限额的员工,从而将 Claude 需要推理的内容从数百 KB 缩减到寥寥几行。
程序化工具调用需要使用工具版本为 code_execution_20260120 或更高版本的代码执行工具。要在发送请求之前检查某个模型是否支持程序化工具调用,请从 Models API 读取其 capabilities.code_execution.supported 值。使用 Models API 介绍了该字段。
快速开始
以下示例中,Claude 以编程方式多次查询数据库并汇总结果。在工具定义中添加 allowed_callers: ["code_execution_20260120"] 即可使该工具能够在代码执行中被调用(请参阅 allowed_callers 字段):
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "Query sales data for the West, East, and Central regions, then tell me which region had the highest revenue",
}
],
tools=[
{"type": "code_execution_20260120", "name": "code_execution"},
{
"name": "query_database",
"description": "Execute a SQL query against the sales database. Returns a list of rows as JSON objects.",
"input_schema": {
"type": "object",
"properties": {
"sql": {"type": "string", "description": "SQL query to execute"}
},
"required": ["sql"],
},
"allowed_callers": ["code_execution_20260120"],
},
],
)
print(response)响应以 stop_reason: "tool_use" 停止,并包含一个 container ID 以及一个针对 query_database 的 tool_use 块,其 caller 字段标识了调用它的代码执行运行。请按照示例工作流的第 3 步所示返回结果,以便代码能够完成执行。
编程式工具调用的工作原理
当您将某个工具配置为可从代码执行中调用,并且 Claude 判断需要该工具时:
- Claude 编写 Python 代码,将该工具作为函数调用,其中可能包含多次工具调用以及前置/后置处理逻辑
- Claude 通过代码执行在沙盒容器中运行此代码
- 当调用工具函数时,代码执行暂停,API 返回一个
tool_use块 - 您提供工具结果,代码执行继续(中间结果不会加载到 Claude 的上下文窗口中)
- 所有代码执行完成后,Claude 接收最终输出并继续处理任务
这种方法特别适用于:
- 大规模数据处理: 在工具结果进入 Claude 的上下文之前对其进行过滤或汇总
- 多步骤工作流: 通过串行或循环调用工具而无需在工具调用之间对 Claude 进行采样,从而节省令牌并降低延迟
- 条件逻辑: 根据中间工具结果做出决策
核心概念
allowed_callers 字段
allowed_callers 字段指定哪些上下文可以调用某个工具:
{
"name": "query_database",
"description": "Execute a SQL query against the database",
"input_schema": {
// ...
},
"allowed_callers": ["code_execution_20260120"]
}可能的取值:
["direct"]- 引导 Claude 直接调用此工具(省略时的默认值)["code_execution_20260120"]- 引导 Claude 仅在代码执行中调用此工具["direct", "code_execution_20260120"]- Claude 可以直接调用此工具,也可以在代码执行中调用
"code_execution_20260120" 和 "code_execution_20260521" 在 allowed_callers 中均被接受且可互换:使用任一代码执行工具版本的请求都能满足列出任一调用方的工具。无论请求声明的是哪个版本,响应块始终将调用方标记为 code_execution_20260120。
响应中的 caller 字段
每个工具使用块都包含一个 caller 字段,指示其调用方式:
直接调用(传统工具使用):
{
"type": "tool_use",
"id": "toolu_abc123",
"name": "query_database",
"input": { "sql": "<sql>" },
"caller": { "type": "direct" }
}编程式调用:
{
"type": "tool_use",
"id": "toolu_xyz789",
"name": "query_database",
"input": { "sql": "<sql>" },
"caller": {
"type": "code_execution_20260120",
"tool_id": "srvtoolu_abc123"
}
}tool_id 是发起调用的代码执行 server_tool_use 块的 id,因此您可以将每个编程式 tool_use 与产生它的代码执行运行相匹配。
容器生命周期
编程式工具调用使用与代码执行相同的容器:
- 容器创建: 除非您复用现有容器,否则每个请求都会创建一个新容器
- 容器 ID: 在响应的
container字段中返回,同时附带expires_at时间戳 - 复用: 在下一个请求中传回容器 ID 以保持状态。当编程式工具调用正在等待您的结果时,该请求中的容器 ID 是必需的而非可选的:没有它,API 会拒绝该请求。
- 过期:
expires_at告诉您容器还剩多少时间。空闲容器目前会在大约 5 分钟后被回收,且任何容器在创建超过 30 天后都无法再被复用。
示例工作流
以下是完整的编程式工具调用流程的工作方式:
第 1 步:初始请求
发送一个包含代码执行和允许编程式调用的工具的请求。要启用编程式调用,请在工具定义中添加 allowed_callers 字段。
请求结构与快速开始示例完全相同:在工具列表中包含 code_execution,为您希望 Claude 从代码中调用的任何工具添加 allowed_callers: ["code_execution_20260120"],然后发送您的用户消息。本工作流的其余步骤使用用户消息 "Query customer purchase history from the last quarter and identify our top 5 customers by revenue"。
第 2 步:包含工具调用的 API 响应
Claude 编写调用您工具的代码。API 暂停并返回:
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "I'll query the purchase history and analyze the results."
},
{
"type": "server_tool_use",
"id": "srvtoolu_abc123",
"name": "code_execution",
"input": {
"code": "import json\n\nrows = json.loads(await query_database({'sql': '<sql>'}))\ntop_customers = sorted(rows, key=lambda x: x['revenue'], reverse=True)[:5]\nprint(f'Top 5 customers: {top_customers}')"
}
},
{
"type": "tool_use",
"id": "toolu_def456",
"name": "query_database",
"input": { "sql": "<sql>" },
"caller": {
"type": "code_execution_20260120",
"tool_id": "srvtoolu_abc123"
}
}
],
"container": {
"id": "container_xyz789",
"expires_at": "2026-01-20T14:30:00Z"
},
"stop_reason": "tool_use"
}第 3 步:提供工具结果
发送完整的对话历史以及您的工具结果。此请求有三个细节需要注意:
- 承载您结果的用户消息只能包含
tool_result块。请参阅消息格式限制。 - 传入暂停响应中的
containerID。如果续接请求存在待处理的编程式工具调用但没有容器 ID,API 会拒绝该请求。 - 发送与原始请求相同的
tools数组。代码执行工具必须仍然存在,暂停的代码才能恢复;并且您在此请求中发送的工具就是 Claude 和正在运行的代码在本轮剩余部分可以使用的定义。
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
container="container_xyz789", # Reuse the container
messages=[
{
"role": "user",
"content": "Query customer purchase history from the last quarter and identify our top 5 customers by revenue",
},
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "I'll query the purchase history and analyze the results.",
},
{
"type": "server_tool_use",
"id": "srvtoolu_abc123",
"name": "code_execution",
"input": {"code": "..."},
},
{
"type": "tool_use",
"id": "toolu_def456",
"name": "query_database",
"input": {"sql": "<sql>"},
"caller": {
"type": "code_execution_20260120",
"tool_id": "srvtoolu_abc123",
},
},
],
},
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_def456",
"content": '[{"customer_id": "C1", "revenue": 45000}, {"customer_id": "C2", "revenue": 38000}, ...]',
}
],
},
],
# 与原始请求相同的 tools 数组
tools=[
{"type": "code_execution_20260120", "name": "code_execution"},
{
"name": "query_database",
"description": "Execute a SQL query against the sales database. Returns a list of rows as JSON objects.",
"input_schema": {
"type": "object",
"properties": {
"sql": {"type": "string", "description": "SQL query to execute"}
},
"required": ["sql"],
},
"allowed_callers": ["code_execution_20260120"],
},
],
)
print(response)第 4 步:下一次工具调用或完成
代码从暂停处继续并处理您的结果。每个续接响应要么再次暂停并返回更多编程式 tool_use 块,要么完成代码执行并让 Claude 继续本轮(第 5 步)。检查 stop_reason 和每个 tool_use 块的 caller 以区分这两种情况:为等待您而暂停的响应具有 stop_reason: "tool_use",以及一个 caller 指明代码执行版本的 tool_use 块;此时您需重复第 3 步,在一条用户消息中为每个待处理的编程式调用提供一个 tool_result。
第 5 步:最终响应
代码执行完成后,Claude 提供最终响应:
{
"content": [
{
"type": "code_execution_tool_result",
"tool_use_id": "srvtoolu_abc123",
"content": {
"type": "code_execution_result",
"stdout": "Top 5 customers: [{'customer_id': 'C1', 'revenue': 45000}, {'customer_id': 'C2', 'revenue': 38000}, {'customer_id': 'C5', 'revenue': 32000}, {'customer_id': 'C8', 'revenue': 28500}, {'customer_id': 'C3', 'revenue': 24000}]",
"stderr": "",
"return_code": 0,
"content": []
}
},
{
"type": "text",
"text": "I've analyzed the purchase history from last quarter. Your top 5 customers generated $167,500 in total revenue, with Customer C1 leading at $45,000."
}
],
"stop_reason": "end_turn"
}高级模式
使用循环进行批处理
Claude 可以编写高效处理多个项目的代码:
regions = ["West", "East", "Central", "North", "South"]
results = {}
for region in regions:
rows = json.loads(await query_database({"sql": f"<sql for {region}>"}))
results[region] = sum(row["revenue"] for row in rows)
# 以编程方式处理结果
top_region = max(results.items(), key=lambda x: x[1])
print(f"Top region: {top_region[0]} with ${top_region[1]:,} in revenue")此模式:
- 将模型往返次数从 N 次(每个区域一次)减少到 1 次
- 在返回给 Claude 之前以编程方式处理大型结果集
- 通过仅返回汇总结论而非原始数据来节省令牌
提前终止
一旦满足成功条件,Claude 即可停止处理:
endpoints = ["us-east", "eu-west", "apac"]
for endpoint in endpoints:
status = await check_health({"endpoint": endpoint})
if status == "healthy":
print(f"Found healthy endpoint: {endpoint}")
break # Stop early, don't check remaining条件式工具选择
path = "/tmp/example.txt"
file_info = json.loads(await get_file_info({"path": path}))
if file_info["size"] < 10000:
content = await read_full_file({"path": path})
else:
content = await read_file_summary({"path": path})
print(content)数据过滤
server_id = "srv-01"
log_text = await fetch_logs({"server_id": server_id})
errors = [line for line in log_text.splitlines() if "ERROR" in line]
print(f"Found {len(errors)} errors")
for error in errors[-10:]: # Only return last 10 errors
print(error)响应格式
编程式工具调用
当代码执行调用工具时:
{
"type": "tool_use",
"id": "toolu_abc123",
"name": "query_database",
"input": { "sql": "<sql>" },
"caller": {
"type": "code_execution_20260120",
"tool_id": "srvtoolu_xyz789"
}
}工具结果处理
您的工具结果会被传回正在运行的代码:
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_abc123",
"content": "[{\"customer_id\": \"C1\", \"revenue\": 45000, \"orders\": 23}, {\"customer_id\": \"C2\", \"revenue\": 38000, \"orders\": 18}, ...]"
}
]
}代码执行完成
当所有工具调用均已满足且代码完成时:
{
"type": "code_execution_tool_result",
"tool_use_id": "srvtoolu_xyz789",
"content": {
"type": "code_execution_result",
"stdout": "Analysis complete. Top 5 customers identified from 847 total records.",
"stderr": "",
"return_code": 0,
"content": []
}
}错误处理
常见错误
| 错误 | 出现位置 | 描述 | 解决方案 |
|---|---|---|---|
invalid_tool_input | 响应中 code_execution_tool_result 错误块上的 error_code | 向代码执行工具传递了无效参数 | 请参阅代码执行工具错误 |
invalid_request_error(针对 tool_choice) | HTTP 400 错误响应 | tool_choice 指定了一个 allowed_callers 不包含 "direct" 的工具 | 要么在该工具的 allowed_callers 中添加 "direct",要么从 tool_choice 中移除该工具并让 Claude 从代码中调用它 |
工具调用期间容器过期
如果您的工具结果未在大约 4 分钟内到达,待处理的调用会在 Claude 正在运行的代码内部抛出 TimeoutError。Claude 会在 stderr 中看到该错误,并通常会重试该调用:
{
"type": "code_execution_tool_result",
"tool_use_id": "srvtoolu_abc123",
"content": {
"type": "code_execution_result",
"stdout": "",
"stderr": "TimeoutError: Calling tool ['query_database'] timed out (no response after 270s).",
"return_code": 0,
"content": []
}
}为防止超时:
- 监控响应中的
expires_at字段 - 为您的工具执行实现超时机制
- 考虑将长时间操作拆分为更小的块
工具执行错误
如果您的工具返回错误:
{
"type": "tool_result",
"tool_use_id": "toolu_abc123",
"content": "Error: Query timeout - table lock exceeded 30 seconds"
}Claude 的代码会收到此错误并可以妥善处理。
约束与限制
功能不兼容
- 结构化输出: 带有
strict: true的工具不支持编程式调用 - 工具选择: 您无法通过
tool_choice强制对特定工具进行编程式调用 - 并行工具使用:
disable_parallel_tool_use: true不支持编程式调用
输入模式限制
input_schema 中包含递归 $ref(引用循环,例如引用自身的模式)的自定义工具无法启用编程式调用。在此类工具的 allowed_callers 中包含代码执行工具版本会导致请求失败,返回 400 invalid_request_error,其消息包含 Circular $ref detected。相同的模式在直接工具调用中是被接受的。
要解决此问题,请执行以下操作之一:
- 通过省略
allowed_callers(或将其设置为["direct"])使该工具仅支持直接调用。同一请求中的其他工具仍可使用编程式调用。 - 从模式中移除循环。例如,将递归展开到固定深度,并在最内层的
description中描述任何更深的嵌套;或者将递归属性替换为一个普通的{"type": "object"},并在其description中说明预期的结构。
工具限制
以下工具无法以编程方式调用:
- 由 MCP 连接器提供的工具
- 计算机使用和浏览器使用工具集(
computer_toolset_20260801和browser_toolset_20260801),其allowed_callers字段仅接受"direct"
消息格式限制
在响应编程式工具调用时,有严格的格式要求:
仅包含工具结果的响应: 如果存在等待结果的待处理编程式工具调用,您的响应消息必须仅包含 tool_result 块。您不能包含任何文本内容,即使是在工具结果之后。
无效 - 响应编程式工具调用时不能包含文本:
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_01",
"content": "[{\"customer_id\": \"C1\", \"revenue\": 45000}]"
},
{ "type": "text", "text": "What should I do next?" }
]
}有效 - 响应编程式工具调用时仅包含工具结果:
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_01",
"content": "[{\"customer_id\": \"C1\", \"revenue\": 45000}]"
}
]
}此限制仅适用于响应编程式(代码执行)工具调用的情况。对于常规的客户端工具调用,您可以在工具结果之后包含文本内容。
仅限文本的工具结果内容: 回应编程式调用的每个 tool_result 的 content 必须是字符串或 text 块。图像、文档和其他内容块类型会被拒绝。
速率限制
编程式工具调用与常规工具调用受相同的速率限制约束。来自代码执行的每次工具调用都计为一次独立调用。
使用前验证工具结果
在实现将以编程方式调用的用户自定义工具时:
- 工具结果以字符串形式返回: 它们可以包含任何内容,包括可能被执行环境处理的代码片段或可执行命令。
- 验证外部工具结果: 如果您的工具返回来自外部来源的数据或接受用户输入,且输出将被解释或作为代码执行,请注意代码注入风险。
令牌效率
编程式工具调用通过三种方式减少令牌消耗:
- 编程式调用的工具结果不会添加到 Claude 的上下文中 - 只有最终的代码输出会被添加
- 中间处理在代码中进行 - 过滤、汇总和其他转换不消耗模型令牌
- 一次代码执行中进行多次工具调用 - 与分开的多个模型轮次相比减少了开销
例如,直接调用 10 个工具所使用的令牌约为以编程方式调用它们并返回摘要的 10 倍。
在 Anthropic 对生产环境 Claude 模型的内部评估中:
- 在一个包含 75 个工具的项目管理智能体基准测试中,启用编程式工具调用使计费输入令牌减少了约 38%,而任务准确率没有变化。
- 在 τ²-bench(航空、零售和电信领域)上,每轮进行一到两次顺序工具调用,编程式工具调用的得分保持不变,而成本增加了约 8%。顺序单次调用的工作流无法从中受益。
- 在生产 API 流量中,
tools数组包含 10 到 49 个工具定义的请求在启用编程式工具调用后通常可节省 20% 到 40% 的令牌。
实际节省量因工作负载形态而异。请参阅何时使用编程式调用。
用量与定价
编程式工具调用采用与代码执行相同的定价。详情请参阅代码执行定价。
最佳实践
工具设计
- 提供详细的输出描述: 由于 Claude 会在代码中反序列化工具结果,请记录其格式(JSON 结构和字段类型)
- 返回结构化数据: JSON 或其他机器可读格式最适合编程式处理
- 保持响应简洁: 仅返回必要的数据以最小化处理开销
何时使用编程式调用
编程式工具调用以少量固定开销(容器启动、脚本生成)换取工具结果令牌和模型往返方面的大幅节省。这种取舍是否划算取决于工作负载形态。
非常适合:
- 跨多个项目的扇出或并行操作(例如,检查 50 个端点或查找 20 条记录)
- 可在进入 Claude 上下文之前进行过滤、汇总或摘要的大型工具结果
- 智能体搜索和检索,其中迭代查询和结果过滤主导整个工作流
不太适合:
- 严格顺序的工作流,其中每次调用都依赖于 Claude 对前一结果的推理,因为在这种情况下脚本无法跳过模型往返
- 少量工具调用且响应较小,尤其是在对话的第一轮,此时容器和脚本开销可能超过节省量
- 需要在调用之间立即获得用户反馈的工具
如果您不确定,请在广泛启用之前,在具有代表性的流量样本上分别测量使用和不使用 allowed_callers 时的计费输入令牌。
性能优化
- 在发出多个相关请求时复用容器以保持状态
- 尽可能在单次代码执行中批量处理相似操作
故障排除
常见问题
设置 tool_choice 时出现 invalid_request_error
tool_choice不能指定allowed_callers中省略了"direct"的工具。要么在该工具的allowed_callers中添加"direct",要么从tool_choice中移除该工具并让 Claude 从代码中调用它。
容器过期
- 请在暂停响应的
expires_at时间戳之前尽早响应每个编程式工具调用。Claude 的代码在大约 4 分钟后停止等待结果,空闲容器目前会在大约 5 分钟后被回收。 - 考虑实现更快的工具执行
工具结果未被正确解析
- 确保您的工具返回 Claude 可以反序列化的字符串数据
- 在工具描述中提供清晰的输出格式文档
调试技巧
- 记录所有工具调用和结果以跟踪流程
- 检查
caller字段以确认编程式调用 - 监控容器 ID 以确保正确复用
- 在启用编程式调用之前独立测试工具
编程式工具调用为何有效
Claude 在大量代码上进行过训练,因此将工具呈现为可调用的 Python 函数可以让它发挥这一优势:
- 工具组合: 链式调用、循环和条件判断是普通的 Python 控制流,而不是一系列模型往返
- 结果处理: Claude 的代码对大型工具输出进行过滤和汇总,或将其写入文件,只有最终输出进入上下文窗口
- 延迟: 在一次代码执行内的工具调用之间不会对模型重新采样
替代实现方案
编程式工具调用是一种可泛化的模式,也可以在您自己的基础设施上实现。以下是各种方法的比较:
客户端直接执行
为 Claude 提供一个代码执行工具,并描述该环境中有哪些可用函数。当 Claude 使用代码调用该工具时,您的应用程序在定义了这些函数的本地环境中执行它。
优点:
- 对应用程序的重构最少
- 完全控制环境和指令
缺点:
- 在沙盒之外执行不受信任的代码
- 工具调用可能成为代码注入的载体
适用场景: 您的应用程序可以安全地执行任意代码,您希望实现最精简,且 Anthropic 的托管方案不符合您的需求。
自行管理的沙盒执行
从 Claude 的角度看方法相同,但代码在具有安全限制(例如,无网络出站)的沙盒容器中运行。如果您的工具需要外部资源,您需要一个在沙盒外执行工具调用的协议。
优点:
- 在您自己的基础设施上安全地进行编程式工具调用
- 完全控制执行环境
缺点:
- 构建和维护复杂
- 需要同时管理基础设施和进程间通信
适用场景: 安全性至关重要,且 Anthropic 的托管解决方案不符合您的要求。
Anthropic 托管执行
Anthropic 的编程式工具调用是沙盒执行的托管版本,配备了为 Claude 调优的、有明确设计取向的 Python 环境。Anthropic 负责容器管理、代码执行和安全的工具调用通信。
优点:
- 默认安全可靠
- 通过工具定义即可启用,无需运行任何基础设施
- 环境和指令针对 Claude 进行了优化
如果您正在使用 Claude API、AWS 上的 Claude Platform 或 Microsoft Foundry,请考虑使用 Anthropic 的托管解决方案。在 Microsoft Foundry 上,编程式工具调用需要托管在 Anthropic 的部署。
数据保留
编程式工具调用构建在代码执行基础设施之上,并使用相同的沙盒容器。容器数据(包括执行产物和输出)最多保留 30 天。
有关所有功能的 ZDR 资格,请参阅 API 与数据保留。
后续步骤
为延迟敏感型应用流式传输工具输入,无需服务器端 JSON 缓冲。
在沙盒容器中运行 Python 和 bash 代码,以分析数据、生成文件并迭代解决方案。
将 Claude 连接到外部工具和 API。了解工具在何处执行、Claude 何时调用它们,以及哪种工具适合您的任务。
指定工具模式、编写有效的描述,并控制 Claude 何时调用您的工具。
Compatibility
- Supported models
- Fable 5 and 5.1
- Mythos 5 and 5.1
- Opus 4.5, 4.6, 4.7, 4.8, 5, and 5.5
- Sonnet 4.5, 4.6, 5, and 5.5
- Haiku 5.5
- Supported platforms
- Claude API
- Claude Platform on AWS
- Microsoft Foundry1
- 在 Microsoft Foundry 上,程序化工具调用需要 Hosted on Anthropic 部署。 ↩
- 程序化工具调用需要使用
code_execution_20260120或更高工具版本的代码执行工具。 - Claude Haiku 4.5 接受
code_execution_20260120及更高的工具版本,但不支持程序化工具调用。
Was this page helpful?