Multi-agent teams under latency pressure and budgets
Our multi-agent team consists of a lead agent with the ability to spawn helper agents to work on parts of a task in parallel. Left to itself, a team optimizes for thoroughness, not speed: nothing in any agent's context says how long the person asking has been waiting, or how long they are willing to wait. In this cookbook, we will elicit Claude's time awareness to accomplish tasks faster.
By the end of this cookbook, you'll be able to:
- Show every agent on a team the same running clock, without breaking prompt caching
- Make a team work faster with one sentence of latency pressure, or pace itself against a latency budget
- Check from a run's own log that every agent saw the clock
We first make Claude aware of how much time has elapsed, with nothing but the Messages API:
- One shared team clock, started when the task starts.
- Every agent sees the clock before every turn. Each time we call the API for any agent,
we append one short line,
[elapsed 252s], to the end of the newest user message (the task prompt on the first turn, the tool results after that). A helper spawned three minutes into the task sees[elapsed 182s]on its very first turn, not0s, because the clock belongs to the team, not to the agent.
We then give Claude a sense of urgency, through one of two approaches:
- Latency pressure: One sentence at the end of the task prompt, and of every helper's brief, says that time matters. That sentence is the only difference from the stock prompts.
- Latency budget: An augmented clock carries a budget alongside the elapsed time,
e.g.
[elapsed 252s / 600s]. No prompt text comes with it: the second number is all any agent is told about the budget.
That is the whole mechanism, and it gives three arms to compare: no clock at all, the clock with pressure, and the clock with a budget.
A note on models. Latency incentives are a newly emerging way to steer Claude. This notebook defaults to
claude-fable-5-1. The approach has not been fully tested on Claude Opus 5 and earlier models, so expect the effect to vary if you changeCOOKBOOK_MODEL.
There is no task built in. You bring the tools, the prompts and the question; a short smoke test at the end shows the clock reaching every agent.
The orchestration (message hub, messaging tools, spawn tools, agent loop) follows the Async multi-agent orchestration(opens in new tab) cookbook; start there if the shape is unfamiliar.
Prerequisites
- Python 3.11 or newer.
- Required knowledge: comfortable reading
async/awaitPython (tasks, events, timeouts), and familiar with the Messages API tool-use loop (a model turn, its tool calls, the tool results you send back). The Async multi-agent orchestration(opens in new tab) cookbook covers the team mechanics this one builds on. - A Claude API key in
ANTHROPIC_API_KEY(or a.envfile). - Some API credit. Each run is a lead plus up to four helpers, each making several calls. We have
capped calls to fifteen per helper and forty for the lead. Running the
notebook top to bottom is one short team run on
claude-fable-5-1at effortxhigh, about a minute. - Headroom on your rate limits. Up to five agents call the API at once, and on a lower tier the SDK's automatic retries show up as extra wall-clock time rather than as errors.
The team clock
A stopwatch and a one-line renderer. There is one clock object per
task and every agent on the team holds a reference to it. That is what makes a late-spawned
helper see the team's elapsed time rather than its own. The optional budget lives on the same
object, so it is the team's too: a helper spawned 182 seconds into a 600-second budget sees
[elapsed 182s / 600s], not a fresh 600 seconds of its own.
The line is rendered in one place, clock_line, in whole seconds, so that with a budget both
numbers are in the same unit and what is left is one subtraction. The square brackets are only
there to set the line apart from the prompt or tool result it follows.
[elapsed 192s] with a 600-second budget: [elapsed 192s / 600s]
[{'role': 'user',
'content': [{'type': 'text',
'text': 'Find the best settings for our process.'},
{'type': 'text', 'text': '[elapsed 192s]'}]}]The hub and the coordination tools
Coordination tools are the client-side ones from the orchestration cookbook: every agent
can send_message / wait_for_message through an in-memory hub, and the lead can
create_subagents. A helper's final answer is delivered to the lead
automatically when the helper finishes. The hub also keeps the run's per-agent tool-call counts
and its log.
Your task tools
This is the part to fill in. TASK_TOOLS is the list of tool definitions every agent gets (built
with the same small tool helper), and make_task_handlers(hub) returns the functions that run
them, one async fn(agent, input) per tool name, created fresh for each run. Both start empty, so
out of the box the team works from the model's own knowledge. The commented template shows the
shape of one tool.
Note: Client-side tools keep the clock fresh. The harness can only add a clock line between API calls. While a response is being sampled on the server, including any server-side tools it runs there, nothing can be injected, so the agent works from the last clock line it saw until the call returns. Each client-side tool call hands control back to the harness, and the next request carries an updated line. The more of an agent's work goes through client-side tools, the more up-to-date its clock stays.
The agent loop, and where the clock goes
One coroutine runs any agent, lead or helper. It is an ordinary tool-use loop with two additions, both marked in the code:
- Right before every API call,
stamp_elapsedappends the team clock line to the newest user message. On the first turn that message is the task prompt, so the agent knows the elapsed time before it does anything. Afterwards the line goes at the end of the last tool result, next to any messages from other agents. Because the stamp happens at send time, time spent waiting on helpers or on a slow tool shows up too. - When a helper finishes, its final text is posted to the lead's inbox. Before that, it reads any message that reached it while it worked, so a redirect from the lead is not lost.
Two API details. The response content (which can include thinking) goes into the history
exactly as returned; only client-side tool_use blocks need a tool_result from us. And the
top-level cache_control turns on
automatic prompt caching(opens in new tab),
which matters for agents whose history fills up with tool results; because the clock line is
appended at the tail, everything before it is normally still a cache hit on the next call.
An agent's tool calls from one turn run concurrently. The loop also writes what each agent was
shown into the run's log ([task], [clock] and [report] entries), which the last section
reads back.
Prompts
TIME_MATTERSis appended once to the lead's task and to every helper's brief. It is the only difference between the stock prompts and the pressured ones. We have experimented with a number of phrases and believe this one offers an acceptable balance between latency and quality.- A budget adds no prompt text. With
latency_budget_secondsset, the clock line gains a second number,[elapsed 252s / 600s], and that is all any agent is told about it. - The system prompts say nothing about time.
run_team(question) runs the stock team with no clock. latency_pressure=True adds the clock and
the sentence; latency_budget_seconds=N adds the clock with the budget on it. They are separate
arms, so asking for both raises an error. A run prints a progress line every PROGRESS_EVERY
seconds and returns the answer, the wall-clock seconds, per-agent tool-call counts and the full
log.
End-to-end test
One short run to check the wiring before you add your own tools. With TASK_TOOLS empty the team
works from the model's own knowledge, and the question below tells the lead to use its helpers so
that there is something to look at. It is a check, not an example task. After the run, three
small deterministic readers print what the run's own log recorded:
print_summaryprints the answer and each agent's tool calls.show_first_promptsprints the first message the lead and the first helper received, with the clock line that followed it. Under pressure, the time-matters sentence is at the end of both.show_clock_linesprints, per agent, the clock lines it was shown. A helper's first line is the team's elapsed time, not0s.show_timelineprints when the lead spawned helpers, when each reported, and when the answer came, against the budget if there is one.
To check the budget arm instead, swap in the commented line.
======================================================================================== ANSWER after 9s wall-clock, 4 helper(s) spawned: **Timsort** — Python's built-in `sorted()` / `list.sort()` uses it. | Algorithm | Average | Worst | Stable | |-----------|---------|-------|--------| | Quicksort | O(n log n) | O(n²) | No | | Mergesort | O(n log n) | O(n log n) | Yes | | Heapsort | O(n log n) | O(n log n) | No | | Timsort | O(n log n) | O(n log n) | Yes | All four helpers reported back; results match standard references. TOTAL task-tool calls 0 --- what the agents were told --- lead, first message: Wiring check. Spawn one helper for each of these four sorting algorithms: quicksort, mergesort, heapsort and timsort. Have each report the average and worst-case time complexity and whether the sort i [...] mall table, and say on the first line which of the four Python's built-in sort uses. Time matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better. [elapsed 0s] helper1, first message: Report, in one short line, for the named sorting algorithm: average-case time complexity, worst-case time complexity, and whether it is stable. Be concise; no extra commentary. Algorithm: quicksort Time matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better. [elapsed 4s] --- the clock lines each agent saw --- lead [elapsed 0s] -> [elapsed 4s] -> [elapsed 5s] -> [elapsed 6s] helper1 [elapsed 4s] helper2 [elapsed 4s] helper3 [elapsed 4s] helper4 [elapsed 4s] --- timeline --- 0s lead starts 4s lead spawns helper1, helper2, helper3, helper4 5s helper1 reports 5s helper4 reports 6s helper3 reports 6s helper2 reports 9s lead answers
Conclusion
We gave a lead-and-helpers team one shared clock and showed it to every agent before every turn. One sentence of latency pressure, or a budget on the clock line, is then what changes how the team paces itself.
- Reach for pressure when sooner is simply better and you have no number in mind. It needs no tuning. In our experience the effect on answer quality is small, but this notebook does not measure it, so check it on your own task.
- Reach for a budget when you have a latency target in mind and want the team to pace itself against it.
- On your own task, keep
TeamClock,clock_line,stamp_elapsedandTIME_MATTERS, and bring your own tools, system prompts and questions. Pick work where thoroughness and speed pull against each other: a task with nothing to trade looks the same in every arm. The swap points here areTASK_TOOLS,make_task_handlers, the two system prompts and your question.
A budget suits longer tasks better than brief ones. When the whole task takes a minute or two, half of that leaves little room to do any of the work, and Claude may overrun the budget rather than abandon the task. Treat the code here as a starting point for your own experiments with latency incentives.
Before acting on a faster arm, check that its answers are still good enough for your task.
The natural next step is to measure: run each arm several times on your own task and compare wall-clock times. For the orchestration itself, see Async multi-agent orchestration(opens in new tab).