Claude Platform Docs
Managed Agents將工作委派給您的代理

定義成果

告訴代理「完成」是什麼樣子,並讓它反覆迭代直到達成目標。

「Outcome」(成果)會告訴工作階段最終結果應該是什麼樣子,以及如何衡量其品質。代理會朝著該目標努力,自我評估並反覆迭代,直到達成成果為止。

當您定義成果時,「harness」(執行框架)會自動配置一個 grader(評分器),依據「rubric」(評分標準)來評估「artifact」(產出物)。評分器使用獨立的「context window」(上下文視窗),以避免受到主要代理實作選擇的影響。

評分器會回傳一段說明,摘要哪些標準通過或未通過,或確認產出物符合評分標準。這些回饋會交回給代理,用於下一次迭代。

建立評分標準

評分標準是一份描述各項標準評分方式的 markdown 文件。評分標準為必填。

評分標準範例:

# DCF Model Rubric

## Revenue Projections
- Uses historical revenue data from the last 5 fiscal years
- Projects revenue for at least 5 years forward
- Growth rate assumptions are explicitly stated and reasonable

## Cost Structure
- COGS and operating expenses are modeled separately
- Margins are consistent with historical trends or deviations are justified

## Discount Rate
- WACC is calculated with stated assumptions for cost of equity and cost of debt
- Beta, risk-free rate, and equity risk premium are sourced or justified

## Terminal Value
- Uses either perpetuity growth or exit multiple method (stated which)
- Terminal growth rate does not exceed long-term GDP growth

## Output Quality
- All figures are in a single .xlsx file with clearly labeled sheets
- Key assumptions are on a separate "Assumptions" sheet
- Sensitivity analysis on WACC and terminal growth rate is included

您可以在 user.define_outcome 上以內嵌文字傳遞評分標準(請參閱建立具有成果的工作階段),或透過 Files API 上傳,以便在多個工作階段之間重複使用。

import time
from pathlib import Path

from anthropic import Anthropic

client = Anthropic()

RUBRIC = """# DCF Model Rubric

## Revenue Projections
- Uses historical revenue data from the last 5 fiscal years
- Projects revenue for at least 5 years forward

## Output Quality
- All figures are in a single .xlsx file with clearly labeled sheets
"""
Path("/tmp/rubric.md").write_text(RUBRIC)

rubric = client.files.upload(file=Path("/tmp/rubric.md"))
print(f"Uploaded rubric: {rubric.id}")

建立具有成果的工作階段

以下範例會為現有的代理和環境(兩者皆另行建立)建立一個工作階段,然後傳送 user.define_outcome 事件。代理會立即開始工作,不需要額外的使用者訊息事件。

# Create a session
session = client.beta.sessions.create(
    agent=agent.id,
    environment_id=environment.id,
    title="Financial analysis on Costco",
)

# Define the outcome — agent starts working on receipt
client.beta.sessions.events.send(
    session_id=session.id,
    events=[
        {
            "type": "user.define_outcome",
            "description": "Build a DCF model for Costco in .xlsx",
            "rubric": {"type": "text", "content": RUBRIC},
            # or: "rubric": {"type": "file", "file_id": rubric.id},
            "max_iterations": 5,  # optional; default 3, max 20
        }
    ],
)

成果事件

以成果為導向的工作階段進度會顯示在事件串流上。

  • agent.* 事件(例如訊息和工具使用)會顯示朝向成果的進度。
  • span.outcome_evaluation_* 事件僅會在以成果為導向的工作階段中發出,並顯示迭代循環的次數以及評分器的回饋過程。
  • 您也可以向以成果為導向的工作階段傳送 user.message 事件,在代理工作進行時引導其方向,但這並非必要:代理會自行朝成果努力,反覆迭代直到成功或用盡迭代次數為止。
  • user.interrupt 事件會暫停目前成果的工作,並將 span.outcome_evaluation_end.result 標記為 interrupted,讓您可以啟動新的成果。
  • 在最後一次成果評估之後,工作階段可以作為對話式工作階段繼續進行,或開始新的成果。工作階段會保留先前成果的歷史記錄。

定義成果使用者事件

這是您用來啟動成果的事件。收到後會回傳此事件,其中包含 processed_at 時間戳記和 outcome_id。

{
  "type": "user.define_outcome",
  "description": "Build a DCF model for Costco in .xlsx",
  "rubric": { "type": "file", "file_id": "file_01..." },
  "max_iterations": 5
}

成果評估開始

當評分器針對一個迭代循環開始評估時發出。iteration 欄位是從 0 開始的修訂計數器:0 是第一次評估,1 是第一次修訂後的重新評估,依此類推。

{
  "type": "span.outcome_evaluation_start",
  "id": "sevt_01def...",
  "outcome_id": "outc_01a...",
  "iteration": 0,
  "processed_at": "2026-03-25T14:01:45Z"
}

成果評估進行中

評分器執行期間發出的心跳訊號。評分器的內部推理是不透明的:您只能看到它正在運作,而無法看到它在想什麼。

{
  "type": "span.outcome_evaluation_ongoing",
  "id": "sevt_01ghi...",
  "outcome_id": "outc_01a...",
  "iteration": 0,
  "processed_at": "2026-03-25T14:02:10Z"
}

成果評估結束

當成果評估週期結束時發出:在評分器完成一次迭代的評估之後,或在成果進行中時工作階段被中斷時。result 欄位表示接下來會發生什麼。

結果後續
satisfied工作階段轉換為 idle。
needs_revision代理開始新的迭代週期。
max_iterations_reached在工作階段轉換為 idle 之前,會進行最後一次確認回合。不會再執行任何評估。
failed工作階段轉換為 idle。當評分標準不適用於交付成果時回傳,例如描述與評分標準互相矛盾時。
interrupted當成果進行中時工作階段被中斷時發出,即使評估尚未開始也是如此。如果在中斷之前沒有觸發 outcome_evaluation_start,則 outcome_evaluation_start_id 為空字串。
{
  "type": "span.outcome_evaluation_end",
  "id": "sevt_01jkl...",
  "outcome_evaluation_start_id": "sevt_01def...",
  "outcome_id": "outc_01a...",
  "result": "satisfied",
  "explanation": "All 12 criteria met: revenue projections use 5 years of historical data, WACC assumptions are stated, sensitivity table is included...",
  "iteration": 0,
  "usage": {
    "input_tokens": 2400,
    "output_tokens": 350,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 1800
  },
  "processed_at": "2026-03-25T14:03:00Z"
}

檢查成果狀態

您可以在事件串流上監聽 span.outcome_evaluation_end,或輪詢 GET /v1/sessions/{session_id} 並讀取 outcome_evaluations[].result。在評估完成之前,result 會回報 pending、running 或 evaluating:

session = client.beta.sessions.retrieve(session.id)

for outcome in session.outcome_evaluations:
    print(f"{outcome.outcome_id}: {outcome.result}")
    # outc_01a...: satisfied

取得交付成果

代理會將輸出檔案寫入沙箱內的 /mnt/session/outputs/。若要取得這些檔案,請透過 Files API 以工作階段 ID 作為 scope_id 列出檔案,然後依 ID 下載。依 scope_id 篩選需要在列出請求上加入 managed-agents-2026-04-01 beta 標頭,因此 SDK 和 CLI 範例會透過 beta 命名空間進行該呼叫,並明確傳遞該標頭。檔案會在代理完成寫入後不久出現在清單中,有時會在工作階段進入閒置狀態後幾秒才出現。如果您預期的檔案尚未列出,請稍候片刻再重新列出;一旦檔案出現在清單中,表示其上傳已完成。

# List files produced by this session
# scope_id filtering requires the managed-agents beta on the files request
files = client.beta.files.list(scope_id=session.id, betas=["managed-agents-2026-04-01"])
for file in files:
    print(file.id, file.filename)

# Download a file
if files.data:
    content = client.files.download(files.data[0].id)
    content.write_to_file("/tmp/output.txt")

後續步驟

在建立工作階段時註冊每位使用者的憑證。

傳送事件、串流回應,並在執行途中中斷或重新導向您的工作階段。

上傳檔案並將其掛載到您的沙箱中,以供讀取和處理。

Was this page helpful?