Claude Platform Docs
Managed Agents将工作委派给智能体

定义结果

告诉智能体'完成'是什么样子,并让它迭代直到达成目标。

"outcome"(结果)告诉会话最终结果应该是什么样子,以及如何衡量其质量。智能体朝着该目标工作,不断进行自我评估和迭代,直到达成结果。

当您定义结果时,"harness"(运行框架)会自动配置一个 grader(评分器),根据 "rubric"(评分标准)评估产出物。评分器使用单独的 "context window"(上下文窗口),以避免受到主智能体实现选择的影响。

评分器会返回一段说明,总结哪些标准通过或未通过,或确认产出物满足评分标准。该反馈会交回给智能体,用于下一次迭代。

创建评分标准

评分标准是一个描述各项标准评分方式的 markdown 文档。评分标准是必需的。

评分标准示例:

# DCF Model Rubric

## Revenue Projections
- Uses historical revenue data from the last 5 fiscal years
- Projects revenue for at least 5 years forward
- Growth rate assumptions are explicitly stated and reasonable

## Cost Structure
- COGS and operating expenses are modeled separately
- Margins are consistent with historical trends or deviations are justified

## Discount Rate
- WACC is calculated with stated assumptions for cost of equity and cost of debt
- Beta, risk-free rate, and equity risk premium are sourced or justified

## Terminal Value
- Uses either perpetuity growth or exit multiple method (stated which)
- Terminal growth rate does not exceed long-term GDP growth

## Output Quality
- All figures are in a single .xlsx file with clearly labeled sheets
- Key assumptions are on a separate "Assumptions" sheet
- Sensitivity analysis on WACC and terminal growth rate is included

您可以在 user.define_outcome 上以内联文本形式传递评分标准(请参阅创建带有结果的会话),也可以通过 Files API 上传,以便在多个会话中复用。

import time
from pathlib import Path

from anthropic import Anthropic

client = Anthropic()

RUBRIC = """# DCF Model Rubric

## Revenue Projections
- Uses historical revenue data from the last 5 fiscal years
- Projects revenue for at least 5 years forward

## Output Quality
- All figures are in a single .xlsx file with clearly labeled sheets
"""
Path("/tmp/rubric.md").write_text(RUBRIC)

rubric = client.files.upload(file=Path("/tmp/rubric.md"))
print(f"Uploaded rubric: {rubric.id}")

创建带有结果的会话

以下示例为现有的智能体和环境(两者均单独创建)创建一个会话,然后发送一个 user.define_outcome 事件。智能体会立即开始工作,无需额外发送用户消息事件。

# Create a session
session = client.beta.sessions.create(
    agent=agent.id,
    environment_id=environment.id,
    title="Financial analysis on Costco",
)

# Define the outcome — agent starts working on receipt
client.beta.sessions.events.send(
    session_id=session.id,
    events=[
        {
            "type": "user.define_outcome",
            "description": "Build a DCF model for Costco in .xlsx",
            "rubric": {"type": "text", "content": RUBRIC},
            # or: "rubric": {"type": "file", "file_id": rubric.id},
            "max_iterations": 5,  # optional; default 3, max 20
        }
    ],
)

结果事件

面向结果的会话的进度会通过事件流呈现。

  • agent.* 事件(例如消息和工具使用)显示朝着结果推进的进度。
  • span.outcome_evaluation_* 事件仅在面向结果的会话中发出,显示迭代循环的次数以及评分器的反馈过程。
  • 您也可以向面向结果的会话发送 user.message 事件,以在智能体工作过程中对其进行引导,但这不是必需的:智能体会自行朝着结果工作,不断迭代,直到成功或用完迭代次数。
  • user.interrupt 事件会暂停当前结果的工作,并将 span.outcome_evaluation_end.result 标记为 interrupted,以便您启动新的结果。
  • 在最终的结果评估之后,会话可以作为对话式会话继续进行,也可以启动新的结果。会话会保留先前结果的历史记录。

定义结果用户事件

这是您为启动结果而发送的事件。收到后会回显该事件,其中包含 processed_at 时间戳和 outcome_id。

{
  "type": "user.define_outcome",
  "description": "Build a DCF model for Costco in .xlsx",
  "rubric": { "type": "file", "file_id": "file_01..." },
  "max_iterations": 5
}

结果评估开始

当评分器针对一个迭代循环开始评估时发出。iteration 字段是一个从 0 开始的修订计数器:0 表示第一次评估,1 表示第一次修订后的重新评估,依此类推。

{
  "type": "span.outcome_evaluation_start",
  "id": "sevt_01def...",
  "outcome_id": "outc_01a...",
  "iteration": 0,
  "processed_at": "2026-03-25T14:01:45Z"
}

结果评估进行中

评分器运行期间发出的心跳事件。评分器的内部推理是不透明的:您只能看到它正在工作,而看不到它在思考什么。

{
  "type": "span.outcome_evaluation_ongoing",
  "id": "sevt_01ghi...",
  "outcome_id": "outc_01a...",
  "iteration": 0,
  "processed_at": "2026-03-25T14:02:10Z"
}

结果评估结束

在一个结果评估周期结束时发出:即评分器完成对一次迭代的评估之后,或在结果处于活动状态时会话被中断时。result 字段指示接下来会发生什么。

结果下一步
satisfied会话转换为 idle。
needs_revision智能体开始新的迭代周期。
max_iterations_reached在会话转换为 idle 之前,会进行最后一轮确认。不再运行任何评估。
failed会话转换为 idle。当评分标准不适用于交付物时返回,例如描述与评分标准相互矛盾时。
interrupted当结果处于活动状态时会话被中断时发出,即使评估尚未开始。如果在中断之前没有触发 outcome_evaluation_start,则 outcome_evaluation_start_id 为空字符串。
{
  "type": "span.outcome_evaluation_end",
  "id": "sevt_01jkl...",
  "outcome_evaluation_start_id": "sevt_01def...",
  "outcome_id": "outc_01a...",
  "result": "satisfied",
  "explanation": "All 12 criteria met: revenue projections use 5 years of historical data, WACC assumptions are stated, sensitivity table is included...",
  "iteration": 0,
  "usage": {
    "input_tokens": 2400,
    "output_tokens": 350,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 1800
  },
  "processed_at": "2026-03-25T14:03:00Z"
}

检查结果状态

您可以在事件流上监听 span.outcome_evaluation_end,也可以轮询 GET /v1/sessions/{session_id} 并读取 outcome_evaluations[].result。在评估完成之前,result 会报告 pending、running 或 evaluating:

session = client.beta.sessions.retrieve(session.id)

for outcome in session.outcome_evaluations:
    print(f"{outcome.outcome_id}: {outcome.result}")
    # outc_01a...: satisfied

获取交付物

智能体会将输出文件写入沙箱内的 /mnt/session/outputs/。要获取这些文件,请通过 Files API 列出文件,并将会话 ID 作为 scope_id,然后按 ID 下载。按 scope_id 筛选需要在列表请求中包含 managed-agents-2026-04-01 beta 标头,因此 SDK 和 CLI 示例通过 beta 命名空间发起该调用,并显式传递该标头。文件会在智能体写入完成后不久出现在列表中,有时会在会话进入空闲状态几秒钟之后才出现。如果您期望的文件尚未列出,请稍等片刻后再次列出;一旦文件出现在列表中,就表示其上传已完成。

# List files produced by this session
# scope_id filtering requires the managed-agents beta on the files request
files = client.beta.files.list(scope_id=session.id, betas=["managed-agents-2026-04-01"])
for file in files:
    print(file.id, file.filename)

# Download a file
if files.data:
    content = client.files.download(files.data[0].id)
    content.write_to_file("/tmp/output.txt")

后续步骤

在创建会话时注册每个用户的凭据。

发送事件、流式传输响应,并在执行过程中中断或重定向您的会话。

上传文件并将其挂载到您的沙箱中,以供读取和处理。

Was this page helpful?