Create a Message Batch
Send a batch of Message creation requests.
The Message Batches API can be used to process multiple Messages API requests at once. Once a Message Batch is created, it begins processing immediately. Batches can take up to 24 hours to complete.
Learn more about the Message Batches API in our user guide
Parameters
params: BatchCreateParams { requests; user_profile_id; workspace_id }
requests: Array<Request>List of requests for prompt completion. Each is an individual request to create a Message.
List of requests for prompt completion. Each is an individual request to create a Message.
custom_id: stringDeveloper-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.
Developer-provided ID created for each request in a Message Batch. Useful for matching results to requests, as results may be given out of request order.
Must be unique for each request within the Message Batch.
params: Params { max_tokens; messages; model; /* 16 more */ }Messages API creation parameters for the individual request.
Messages API creation parameters for the individual request.
See the Messages API reference for full documentation on available parameters.
max_tokens: numberThe maximum number of tokens to generate before stopping.
The maximum number of tokens to generate before stopping.
Note that our models may stop before reaching this maximum. This parameter only specifies the absolute maximum number of tokens to generate.
Set to 0 to populate the prompt cache without generating a response.
Different models have different maximum values for this parameter. See models for details.
messages: Array<MessageParam { content; role }>Input messages.
Input messages.
Our models are trained to operate on alternating user and assistant conversational turns. When creating a new Message, you specify the prior conversational turns with the messages parameter, and the model then generates the next Message in the conversation. Consecutive user or assistant turns in your request will be combined into a single turn.
Each input message must be an object with a role and content. You can specify a single user-role message, or you can include multiple user and assistant messages.
If the final message uses the assistant role, the response content will continue immediately from the content in that message. This can be used to constrain part of the model's response. This is called prefill. On models that don't support prefill, creating a message that ends with a partial assistant response returns a 400 error. See Prefill not supported.
Example with a single user message:
[{"role": "user", "content": "Hello, Claude"}]Example with multiple conversational turns:
[
{"role": "user", "content": "Hello there."},
{"role": "assistant", "content": "Hi, I'm Claude. How can I help you?"},
{"role": "user", "content": "Can you explain LLMs in plain English?"},
]Example with a partially-filled response from Claude, for models that support prefill:
[
{"role": "user", "content": "What's the Greek name for Sun? (A) Sol (B) Helios (C) Sun"},
{"role": "assistant", "content": "The best answer is ("},
]Each input message content may be either a single string or an array of content blocks, where each block has a specific type. Using a string for content is shorthand for an array of one content block of type "text". The following input messages are equivalent:
{"role": "user", "content": "Hello, Claude"}{"role": "user", "content": [{"type": "text", "text": "Hello, Claude"}]}See input examples.
Note that if you want to include a system prompt, you can use the top-level system parameter — there is no "system" role for input messages in the Messages API.
There is a limit of 100,000 messages in a single request.
content: string | Array<ContentBlockParam>
Array<ContentBlockParam>
interface TextBlockParam { type: "text"; text; cache_control; citations }
interface ImageBlockParam { type: "image"; source; cache_control; transformations }
interface DocumentBlockParam { type: "document"; source; cache_control; /* 3 more */ }
interface SearchResultBlockParam { type: "search_result"; content; source; /* 3 more */ }
interface ThinkingBlockParam { type: "thinking"; signature; thinking }
signature: stringThe signature value of this thinking block, exactly as returned by the API in a previous response. Used to verify that the block was generated by Claude.
The signature value of this thinking block, exactly as returned by the API in a previous response. Used to verify that the block was generated by Claude.
Thinking blocks must be passed back unmodified and in their original order; a modified block results in a 400 invalid_request_error.
The thinking text of this block as returned by the API.
interface RedactedThinkingBlockParam { type: "redacted_thinking"; data }
The data value of this redacted thinking block, exactly as returned by the API in a previous response. Opaque and encrypted; pass it back unchanged.
interface ToolUseBlockParam { type: "tool_use"; id; input; /* 4 more */ }
interface ToolResultBlockParam { type: "tool_result"; tool_use_id; cache_control; /* 3 more */ }
interface ServerToolUseBlockParam { type: "server_tool_use"; id; input; /* 3 more */ }
interface WebSearchToolResultBlockParam { type: "web_search_tool_result"; content; tool_use_id; /* 2 more */ }
interface WebFetchToolResultBlockParam { type: "web_fetch_tool_result"; content; tool_use_id; /* 2 more */ }
interface CodeExecutionToolResultBlockParam { type: "code_execution_tool_result"; content; tool_use_id; cache_control }
tool_use_id: string
cache_control?: CacheControlEphemeral | nullCreate a cache control breakpoint at this content block.
Create a cache control breakpoint at this content block.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
interface BashCodeExecutionToolResultBlockParam { type: "bash_code_execution_tool_result"; content; tool_use_id; cache_control }
interface TextEditorCodeExecutionToolResultBlockParam { type: "text_editor_code_execution_tool_result"; content; tool_use_id; cache_control }
interface ToolSearchToolResultBlockParam { type: "tool_search_tool_result"; content; tool_use_id; cache_control }
interface ContainerUploadBlockParam { type: "container_upload"; file_id; cache_control }A content block that represents a file to be uploaded to the container
Files uploaded via this block will be available in the container's input directory.
A content block that represents a file to be uploaded to the container Files uploaded via this block will be available in the container's input directory.
cache_control?: CacheControlEphemeral | nullCreate a cache control breakpoint at this content block.
Create a cache control breakpoint at this content block.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
role: "user" | "assistant" | "system"
model: ModelThe model that will complete your prompt.
The model that will complete your prompt.
See models for additional details and options.
Fastest model for high-volume, real-time tasks
Efficient model for coding and agents
Frontier intelligence for ambitious tasks across coding, scientific discovery, and enterprise workflows
Powerful intelligence for coding, knowledge work, and long-running agents
Our most capable model for cybersecurity and biology research, available through trusted access programs
Efficient model for coding and agents
Next generation of intelligence for the hardest knowledge work and coding problems
Most capable model for cybersecurity and biology research
Powerful intelligence for long-running agents and coding
Powerful intelligence for long-running agents and coding
Powerful intelligence for long-running agents and coding
Powerful intelligence for long-running agents and coding
Best combination of speed and intelligence
Powerful intelligence for long-running agents and coding
Powerful intelligence for long-running agents and coding
"claude-mythos-preview"DeprecatedNew class of intelligence, strongest in coding and cybersecurity
New class of intelligence, strongest in coding and cybersecurity
"claude-sonnet-4-5"DeprecatedHigh-performance model for agents and coding
High-performance model for agents and coding
"claude-sonnet-4-5-20250929"DeprecatedHigh-performance model for agents and coding
High-performance model for agents and coding
cache_control?: CacheControlEphemeral | nullTop-level cache control automatically applies a cache_control marker to the last cacheable block in the request.
Top-level cache control automatically applies a cache_control marker to the last cacheable block in the request.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
container?: MessageCreateParamsContainer | nullContainer identifier for reuse across requests.
Container identifier for reuse across requests.
interface ContainerParams { id; skills }Container parameters with skills to be loaded.
Container parameters with skills to be loaded.
Container id
skills?: Array<SkillParams> | nullList of skills to load in the container
List of skills to load in the container
type: "anthropic" | "custom"Type of skill - either 'anthropic' (built-in) or 'custom' (user-defined)
Type of skill - either 'anthropic' (built-in) or 'custom' (user-defined)
skill_id: stringSkill ID
Skill ID
version?: stringSkill version or 'latest' for most recent version
Skill version or 'latest' for most recent version
diagnostics?: DiagnosticsParam | nullRequest-level diagnostics. Supply previous_message_id to have the response include diagnostics.cache_miss_reason explaining any prompt-cache divergence from that prior request.
Request-level diagnostics. Supply previous_message_id to have the response include diagnostics.cache_miss_reason explaining any prompt-cache divergence from that prior request.
previous_message_id?: string | nullThe id (msg_...) from this client's previous /v1/messages response. The server compares that request's prompt fingerprint against this one and returns diagnostics.cache_miss_reason when the prompt-cache prefix could not be reused. Pass null on the first turn to opt in without a prior message to compare.
The id (msg_...) from this client's previous /v1/messages response. The server compares that request's prompt fingerprint against this one and returns diagnostics.cache_miss_reason when the prompt-cache prefix could not be reused. Pass null on the first turn to opt in without a prior message to compare.
Specifies the geographic region for inference processing. If not specified, the workspace's default_inference_geo is used.
metadata?: Metadata { user_id }An object describing metadata about the request.
An object describing metadata about the request.
user_id?: string | nullAn external identifier for the user who is associated with the request.
An external identifier for the user who is associated with the request.
This should be a uuid, hash value, or other opaque identifier. Anthropic may use this id to help detect abuse. Do not include any identifying information such as name, email address, or phone number.
output_config?: OutputConfig { effort; format }Configuration options for the model's output, such as the output format.
Configuration options for the model's output, such as the output format.
effort?: "low" | "medium" | "high" | 2 more | nullHow much effort the model should put into its response. Higher effort levels may result in more thorough analysis but take longer.
How much effort the model should put into its response. Higher effort levels may result in more thorough analysis but take longer.
Valid values are low, medium, high, xhigh, or max.
format?: JSONOutputFormat | nullA schema to specify Claude's output format in responses. See structured outputs
A schema to specify Claude's output format in responses. See structured outputs
The JSON schema of the format
service_tier?: "auto" | "standard_only"Determines whether to use priority capacity (if available) or standard capacity for this request.
Determines whether to use priority capacity (if available) or standard capacity for this request.
Anthropic offers different levels of service for your API requests. See service-tiers for details.
stop_sequences?: Array<string>Custom text sequences that will cause the model to stop generating.
Custom text sequences that will cause the model to stop generating.
Our models will normally stop when they have naturally completed their turn, which will result in a response stop_reason of "end_turn".
If you want the model to stop generating when it encounters custom strings of text, you can use the stop_sequences parameter. If the model encounters one of the custom sequences, the response stop_reason value will be "stop_sequence" and the response stop_sequence value will contain the matched stop sequence.
stream?: booleanWhether to incrementally stream the response using server-sent events. When true, SDKs return a raw event stream.
Whether to incrementally stream the response using server-sent events. When true, SDKs return a raw event stream.
In the TypeScript, Python and Ruby SDKs, the recommended way to stream is messages.stream(). It sets stream for you and accumulates the events into the final message. See Streaming with SDKs for an example in each language.
system?: string | Array<TextBlockParam>System prompt.
System prompt.
A system prompt is a way of providing context and instructions to Claude, such as specifying a particular goal or role. See our guide to system prompts.
Array<TextBlockParam { type: "text"; text; cache_control; citations }>
text: string
cache_control?: CacheControlEphemeral | nullCreate a cache control breakpoint at this content block.
Create a cache control breakpoint at this content block.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
citations?: Array<TextCitationParam> | null
interface CitationCharLocationParam { type: "char_location"; cited_text; document_index; /* 3 more */ }
document_index: number
document_title: string | null
start_char_index: number
interface CitationPageLocationParam { type: "page_location"; cited_text; document_index; /* 3 more */ }
document_index: number
document_title: string | null
start_page_number: number
interface CitationContentBlockLocationParam { type: "content_block_location"; cited_text; document_index; /* 3 more */ }
cited_text: stringThe full text of the cited block range, concatenated.
The full text of the cited block range, concatenated.
Always equals the contents of content[start_block_index:end_block_index] joined together. The text block is the minimal citable unit; this field is never a substring of a single block. Not counted toward output tokens, and not counted toward input tokens when sent back in subsequent turns.
document_index: number
document_title: string | null
end_block_index: numberExclusive 0-based end index of the cited block range in the source's content array.
Exclusive 0-based end index of the cited block range in the source's content array.
Always greater than start_block_index; a single-block citation has end_block_index = start_block_index + 1.
start_block_index: number0-based index of the first cited block in the source's content array.
0-based index of the first cited block in the source's content array.
interface CitationWebSearchResultLocationParam { type: "web_search_result_location"; cited_text; encrypted_index; /* 2 more */ }
title: string | null
url: string
interface CitationSearchResultLocationParam { type: "search_result_location"; cited_text; end_block_index; /* 4 more */ }
cited_text: stringThe full text of the cited block range, concatenated.
The full text of the cited block range, concatenated.
Always equals the contents of content[start_block_index:end_block_index] joined together. The text block is the minimal citable unit; this field is never a substring of a single block. Not counted toward output tokens, and not counted toward input tokens when sent back in subsequent turns.
end_block_index: numberExclusive 0-based end index of the cited block range in the source's content array.
Exclusive 0-based end index of the cited block range in the source's content array.
Always greater than start_block_index; a single-block citation has end_block_index = start_block_index + 1.
search_result_index: number0-based index of the cited search result among all search_result content blocks in the request, in the order they appear across messages and tool results.
0-based index of the cited search result among all search_result content blocks in the request, in the order they appear across messages and tool results.
Counted separately from document_index; server-side web search results are not included in this count.
start_block_index: number0-based index of the first cited block in the source's content array.
0-based index of the first cited block in the source's content array.
thinking?: ThinkingConfigParamConfiguration for Claude's thinking.
Configuration for Claude's thinking.
With {"type": "adaptive"}, Claude decides when and how much to think. With {"type": "enabled"} (manual extended thinking), you set a budget_tokens of at least 1,024. Thinking tokens count toward your max_tokens limit.
Which type values are accepted, and what happens when you omit thinking, depend on the model. See thinking for each model's behavior.
tool_choice?: ToolChoiceHow the model should use the provided tools. The model can use a specific tool, any available tool, decide by itself, or not use tools at all.
How the model should use the provided tools. The model can use a specific tool, any available tool, decide by itself, or not use tools at all.
tools?: Array<ToolUnion>Definitions of tools that the model may use.
Definitions of tools that the model may use.
If you include tools in your API request, the model may return tool_use content blocks that represent the model's use of those tools. You can then run those tools using the tool input generated by the model and then optionally return results back to the model using tool_result content blocks.
There are two types of tools: client tools and server tools. The behavior described below applies to client tools. For server tools, see their individual documentation as each has its own behavior (e.g., the web search tool).
Each tool definition includes:
name: Name of the tool.description: Optional, but strongly-recommended description of the tool.input_schema: JSON schema for the toolinputshape that the model will produce intool_useoutput content blocks.
For example, if you defined tools as:
[
{
"name": "get_stock_price",
"description": "Get the current stock price for a given ticker symbol.",
"input_schema": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "The stock ticker symbol, e.g. AAPL for Apple Inc."
}
},
"required": ["ticker"]
}
}
]And then asked the model "What's the S&P 500 at today?", the model might produce tool_use content blocks in the response like this:
[
{
"type": "tool_use",
"id": "toolu_01D7FLrfh4GYq7yT1ULFeyMV",
"name": "get_stock_price",
"input": { "ticker": "^GSPC" }
}
]You might then run your get_stock_price tool with {"ticker": "^GSPC"} as an input, and return the following back to the model in a subsequent user message:
[
{
"type": "tool_result",
"tool_use_id": "toolu_01D7FLrfh4GYq7yT1ULFeyMV",
"content": "259.75 USD"
}
]Tools can be used for workflows that include running client-side tools and functions, or more generally whenever you want the model to produce a particular JSON structure of output.
See our guide for more details.
interface Tool { type; input_schema; name; /* 7 more */ }
interface ToolBash20250124 { type: "bash_20250124"; name; allowed_callers; /* 4 more */ }
interface CodeExecutionTool20250522 { type: "code_execution_20250522"; name; allowed_callers; /* 3 more */ }
interface CodeExecutionTool20250825 { type: "code_execution_20250825"; name; allowed_callers; /* 3 more */ }
interface CodeExecutionTool20260120 { type: "code_execution_20260120"; name; allowed_callers; /* 3 more */ }Code execution tool with REPL state persistence (daemon mode + gVisor checkpoint).
Code execution tool with REPL state persistence (daemon mode + gVisor checkpoint).
interface CodeExecutionTool20260521 { type: "code_execution_20260521"; name; allowed_callers; /* 3 more */ }Code execution tool with REPL state persistence.
Code execution tool with REPL state persistence.
interface BrowserToolset20260801 { type: "browser_toolset_20260801"; cache_control; configs }The browser toolset: a single tools[] entry (carrying no
name) that declares the browser tool family. The model is served
the family's tool with any members disabled via configs removed
from its schema.
The browser toolset: a single tools[] entry (carrying no
name) that declares the browser tool family. The model is served
the family's tool with any members disabled via configs removed
from its schema.
cache_control?: CacheControlEphemeral | nullCreate a cache control breakpoint at this content block.
Create a cache control breakpoint at this content block.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
configs?: BrowserToolsetConfigs | nullSparse per-member overrides, keyed by member name. Absent, null, and {} are equivalent; a member's defaults apply wherever its key is absent.
Sparse per-member overrides, keyed by member name. Absent, null, and {} are equivalent; a member's defaults apply wherever its key is absent.
interface MemoryTool20250818 { type: "memory_20250818"; name; allowed_callers; /* 4 more */ }
interface ComputerToolset20260801 { type: "computer_toolset_20260801"; cache_control; configs }The computer toolset: a single tools[] entry (carrying no
name) that declares the computer tool family. The model is
served the family's tool with any members disabled via configs
removed from its schema. Every member is enabled by default, zoom
included. The single-tool options display_number and
enable_zoom are not fields of a toolset entry — it carries only
type, configs, and cache_control; zoom is controlled
via configs.zoom.enabled.
The computer toolset: a single tools[] entry (carrying no
name) that declares the computer tool family. The model is
served the family's tool with any members disabled via configs
removed from its schema. Every member is enabled by default, zoom
included. The single-tool options display_number and
enable_zoom are not fields of a toolset entry — it carries only
type, configs, and cache_control; zoom is controlled
via configs.zoom.enabled.
cache_control?: CacheControlEphemeral | nullCreate a cache control breakpoint at this content block.
Create a cache control breakpoint at this content block.
ttl?: "5m" | "1h"The time-to-live for the cache control breakpoint.
The time-to-live for the cache control breakpoint.
This may be one the following values:
5m: 5 minutes1h: 1 hour
Defaults to 5m. See prompt caching pricing for details.
configs?: ComputerToolsetConfigs | nullSparse per-member overrides, keyed by member name. Absent, null, and {} are equivalent; a member's defaults apply wherever its key is absent.
Sparse per-member overrides, keyed by member name. Absent, null, and {} are equivalent; a member's defaults apply wherever its key is absent.
interface ToolTextEditor20250124 { type: "text_editor_20250124"; name; allowed_callers; /* 4 more */ }
interface ToolTextEditor20250429 { type: "text_editor_20250429"; name; allowed_callers; /* 4 more */ }
interface ToolTextEditor20250728 { type: "text_editor_20250728"; name; allowed_callers; /* 5 more */ }
interface WebSearchTool20250305 { type: "web_search_20250305"; name; allowed_callers; /* 7 more */ }
interface WebFetchTool20250910 { type: "web_fetch_20250910"; name; allowed_callers; /* 9 more */ }
interface WebSearchTool20260209 { type: "web_search_20260209"; name; allowed_callers; /* 7 more */ }
interface WebFetchTool20260209 { type: "web_fetch_20260209"; name; allowed_callers; /* 9 more */ }
interface WebFetchTool20260309 { type: "web_fetch_20260309"; name; allowed_callers; /* 10 more */ }Web fetch tool with use_cache parameter for bypassing cached content.
Web fetch tool with use_cache parameter for bypassing cached content.
interface WebSearchTool20260318 { type: "web_search_20260318"; name; allowed_callers; /* 8 more */ }
interface WebFetchTool20260318 { type: "web_fetch_20260318"; name; allowed_callers; /* 11 more */ }
interface ToolSearchToolBm25_20251119 { type; name; allowed_callers; /* 3 more */ }
interface ToolSearchToolRegex20251119 { type; name; allowed_callers; /* 3 more */ }
temperature?: numberDeprecatedAmount of randomness injected into the response.
Amount of randomness injected into the response.
Defaults to 1.0. Ranges from 0.0 to 1.0. Use temperature closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks.
Note that even with temperature of 0.0, the results will not be fully deterministic.
top_k?: numberDeprecatedOnly sample from the top K options for each subsequent token.
Only sample from the top K options for each subsequent token.
Used to remove "long tail" low probability responses. Learn more technical details here.
Recommended for advanced use cases only.
top_p?: numberDeprecatedUse nucleus sampling.
Use nucleus sampling.
In nucleus sampling, we compute the cumulative distribution over all the options for each subsequent token in decreasing probability order and cut it off once it reaches a particular probability specified by top_p.
Recommended for advanced use cases only.
The user profile ID to attribute the requests in this batch to. Use when acting on behalf of a party other than your organization. Requires the user-profiles beta header. Applies to every request in the batch; an individual request whose user_profile_id body field conflicts with this header is errored.
workspace_id?: stringHeader parameterOptional header to select the Workspace for this request. The value is a Workspace ID (for example, wrkspc_011CZkZaBF1tNoB5wlCeusgy).
Optional header to select the Workspace for this request. The value is a Workspace ID (for example, wrkspc_011CZkZaBF1tNoB5wlCeusgy).
Only needed for credentials that can act on more than one Workspace. A credential that belongs to a specific Workspace may omit it; if sent, it must match that Workspace.
Returns
interface MessageBatch { type: "message_batch"; id; archived_at; /* 7 more */ }
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env["ANTHROPIC_API_KEY"] // This is the default and can be omitted
});
const messageBatch = await client.messages.batches.create({
requests: [
{
custom_id: "my-custom-id-1",
params: {
max_tokens: 1024,
messages: [{ content: "Hello, world", role: "user" }],
model: "claude-opus-5"
}
}
]
});
console.log(messageBatch.id);{
"id": "msgbatch_013Zva2CMHLNnXjNJJKqJ2EF",
"archived_at": "2024-08-20T18:37:24.100435Z",
"cancel_initiated_at": "2024-08-20T18:37:24.100435Z",
"created_at": "2024-08-20T18:37:24.100435Z",
"ended_at": "2024-08-20T18:37:24.100435Z",
"expires_at": "2024-08-20T18:37:24.100435Z",
"processing_status": "in_progress",
"request_counts": {
"canceled": 10,
"errored": 30,
"expired": 10,
"processing": 100,
"succeeded": 50
},
"results_url": "https://api.anthropic.com/v1/messages/batches/msgbatch_013Zva2CMHLNnXjNJJKqJ2EF/results",
"type": "message_batch"
}Returns Examples
{
"id": "msgbatch_013Zva2CMHLNnXjNJJKqJ2EF",
"archived_at": "2024-08-20T18:37:24.100435Z",
"cancel_initiated_at": "2024-08-20T18:37:24.100435Z",
"created_at": "2024-08-20T18:37:24.100435Z",
"ended_at": "2024-08-20T18:37:24.100435Z",
"expires_at": "2024-08-20T18:37:24.100435Z",
"processing_status": "in_progress",
"request_counts": {
"canceled": 10,
"errored": 30,
"expired": 10,
"processing": 100,
"succeeded": 50
},
"results_url": "https://api.anthropic.com/v1/messages/batches/msgbatch_013Zva2CMHLNnXjNJJKqJ2EF/results",
"type": "message_batch"
}