Claude Message Batches API: Cut Your AI Bill in Half
The Claude Message Batches API processes requests in bulk and halves token costs. Here is how it works and when to use it.

Contents 8
The Claude Message Batches API is an official Anthropic service that processes your Claude requests in bulk and asynchronously, cutting all token costs by 50 percent. Same model, same quality, same answers; you just get them in minutes instead of seconds. Here is how it works, when it pays off, and how we use it ourselves.
In short:
- The Message Batches API processes Claude requests in bulk and asynchronously; all token usage is 50 percent off.
- A single batch holds up to 100,000 requests or 256 MB.
- Most batches finish in under an hour; the upper limit is 24 hours.
- Results can be downloaded for 29 days.
- A perfect fit for any job that doesn't need an instant answer: translation, classification, summarization, evaluation.
What is the Claude Message Batches API?
Normally you send Claude a request and expect the answer within seconds. With the Batches API you send hundreds or even tens of thousands of requests at once as a single "batch". Anthropic processes them in the background as capacity allows, and you collect the results together when they are done (Anthropic, Batch processing).
The key point: every request inside a batch is exactly the same as a regular Messages API request. Same model, same system prompt, same tools. Vision, tool use (including web search and code execution), multi-turn conversations and extended thinking all work inside a batch. You give up nothing in quality; you only accept waiting.
What is a 50 percent discount really worth?
The batch price is exactly half the standard price. For current models (per million tokens):
| Model | Standard input / output | Batch input / output |
|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $2 / $10 |
| Claude Sonnet 5.5 | $2 / $10 | $1 / $5 |
| Claude Haiku 4.5 | $1 / $5 | $0.50 / $2.50 |
A concrete example: say you generate descriptions for the 10,000 products of an online store. Each request uses about 1,000 input tokens and 500 output tokens, for a total of 10 million input and 5 million output tokens.
- Claude Sonnet 5.5, standard API: 10 × $2 + 5 × $10 = $70
- Claude Sonnet 5.5, Batches API: 10 × $1 + 5 × $5 = $35
Same texts, same quality, half the price. If the job repeats every month, the difference becomes a real budget line by the end of the year.
When should you use the Batches API, and when not?
The rule is simple: if a user is waiting at the screen, use the standard API; if nobody is waiting, use the Batches API.
Ideal batch jobs:
- Bulk content generation: product descriptions, meta descriptions, e-mail variants.
- Translation: blog posts, app strings, documentation.
- Classification and moderation: sorting customer reviews into positive and negative, spam detection.
- Data analysis: summarizing thousands of support tickets, extracting data from reports.
- Evaluation: measuring how a prompt change affects thousands of test cases.
Jobs that don't fit a batch:
- Live chatbots and customer support assistants.
- Any feature that must show an answer on screen within seconds.
- Interfaces that need streaming.
stream: trueis not supported inside a batch.
Step by step: your first batch in Python
With the official Anthropic Python SDK, creating a batch takes a few lines. Each request gets a unique custom_id; you will use it to match the results.
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
client = anthropic.Anthropic()
reviews = [
"Great product, and shipping was really fast!",
"The package arrived crushed, never buying again.",
"Not bad, fine for the price.",
]
batch = client.messages.batches.create(
requests=[
Request(
custom_id=f"review-{i}",
params=MessageCreateParamsNonStreaming(
model="claude-sonnet-5-5",
max_tokens=50,
messages=[{
"role": "user",
"content": f"Classify this review as positive, negative or neutral, answer in one word: {review}",
}],
),
)
for i, review in enumerate(reviews)
]
)
print(batch.id, batch.processing_status) # msgbatch_..., in_progress
When the batch is created its status is in_progress. You check the status periodically; once it is ended, the results are ready:
import time
while client.messages.batches.retrieve(batch.id).processing_status != "ended":
time.sleep(60)
for result in client.messages.batches.results(batch.id):
if result.result.type == "succeeded":
message = result.result.message
print(result.custom_id, message.content[0].text)
else:
print(result.custom_id, "failed:", result.result.type)
Each result has one of four types: succeeded, errored, canceled or expired. You only need to resend the ones that failed.
Lesser-known details
Where first-time Batches API users usually get stuck:
- Results come back unordered. The result list may not follow the order you sent. Always match by
custom_id, never by position. - Validation happens later. An invalid parameter doesn't fail when you create the batch; it comes back as
erroredafter processing ends. Before sending a large batch, test the request shape with a single standard API call. custom_idrules are strict: 1 to 64 characters, only letters, digits, hyphens and underscores.- A few parameters aren't supported:
stream: true, fast mode (speed) andmax_tokens: 0return a validation error inside a batch. - Combine it with prompt caching. If every request shares the same long system prompt or document, cache it; the caching and batch discounts stack. Because batches can run longer than 5 minutes, Anthropic recommends the 1-hour cache duration.
- Batches are scoped to a workspace and can be canceled. If you sent the wrong job, stop it with
cancel.
How we use the Batches API at EngerekTech
The blog you are reading is published in both Turkish and English. The English translation feature in our portal sends its translation requests through the Batches API instead of the standard API. Getting a translation in a few minutes instead of a few seconds changes nothing for us, but we save half the cost on every translation.
Decisions like this look small one by one. But in a business that builds AI into its product, deciding up front which requests go out instantly and which go in bulk directly shapes the monthly bill. If you are planning to integrate AI into your own product, get in touch with us; let's build an architecture that accounts for cost from day one.
Frequently asked questions
Are batch results lower quality than the standard API?
No. The same model runs with the same parameters. The only difference is when the answer arrives.
How long can a batch take at most?
Most batches finish in under an hour. Requests that can't be completed within 24 hours are marked expired, and you are not billed for them.
How long can I keep the results?
Results can be downloaded for 29 days after the batch is created. Save anything you want to keep permanently in your own database.
Which models are supported?
All of Anthropic's active models work with the Batches API, including Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 4.5.


