All posts

#claude#anthropic#ai

Claude Message Batches API: Cut Your AI Bill in Half

The Claude Message Batches API processes requests in bulk and halves token costs. Here is how it works and when to use it.

Claude Message Batches API: Cut Your AI Bill in Half
Contents 8

The Claude Message Batches API is an official Anthropic service that processes your Claude requests in bulk and asynchronously, cutting all token costs by 50 percent. Same model, same quality, same answers; you just get them in minutes instead of seconds. Here is how it works, when it pays off, and how we use it ourselves.

In short:

  • The Message Batches API processes Claude requests in bulk and asynchronously; all token usage is 50 percent off.
  • A single batch holds up to 100,000 requests or 256 MB.
  • Most batches finish in under an hour; the upper limit is 24 hours.
  • Results can be downloaded for 29 days.
  • A perfect fit for any job that doesn't need an instant answer: translation, classification, summarization, evaluation.

What is the Claude Message Batches API?

Normally you send Claude a request and expect the answer within seconds. With the Batches API you send hundreds or even tens of thousands of requests at once as a single "batch". Anthropic processes them in the background as capacity allows, and you collect the results together when they are done (Anthropic, Batch processing).

The key point: every request inside a batch is exactly the same as a regular Messages API request. Same model, same system prompt, same tools. Vision, tool use (including web search and code execution), multi-turn conversations and extended thinking all work inside a batch. You give up nothing in quality; you only accept waiting.

What is a 50 percent discount really worth?

The batch price is exactly half the standard price. For current models (per million tokens):

Model Standard input / output Batch input / output
Claude Opus 5.5 $4 / $20 $2 / $10
Claude Sonnet 5.5 $2 / $10 $1 / $5
Claude Haiku 4.5 $1 / $5 $0.50 / $2.50

A concrete example: say you generate descriptions for the 10,000 products of an online store. Each request uses about 1,000 input tokens and 500 output tokens, for a total of 10 million input and 5 million output tokens.

  • Claude Sonnet 5.5, standard API: 10 × $2 + 5 × $10 = $70
  • Claude Sonnet 5.5, Batches API: 10 × $1 + 5 × $5 = $35

Same texts, same quality, half the price. If the job repeats every month, the difference becomes a real budget line by the end of the year.

When should you use the Batches API, and when not?

The rule is simple: if a user is waiting at the screen, use the standard API; if nobody is waiting, use the Batches API.

Ideal batch jobs:

  • Bulk content generation: product descriptions, meta descriptions, e-mail variants.
  • Translation: blog posts, app strings, documentation.
  • Classification and moderation: sorting customer reviews into positive and negative, spam detection.
  • Data analysis: summarizing thousands of support tickets, extracting data from reports.
  • Evaluation: measuring how a prompt change affects thousands of test cases.

Jobs that don't fit a batch:

  • Live chatbots and customer support assistants.
  • Any feature that must show an answer on screen within seconds.
  • Interfaces that need streaming. stream: true is not supported inside a batch.

Step by step: your first batch in Python

With the official Anthropic Python SDK, creating a batch takes a few lines. Each request gets a unique custom_id; you will use it to match the results.

import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request

client = anthropic.Anthropic()

reviews = [
    "Great product, and shipping was really fast!",
    "The package arrived crushed, never buying again.",
    "Not bad, fine for the price.",
]

batch = client.messages.batches.create(
    requests=[
        Request(
            custom_id=f"review-{i}",
            params=MessageCreateParamsNonStreaming(
                model="claude-sonnet-5-5",
                max_tokens=50,
                messages=[{
                    "role": "user",
                    "content": f"Classify this review as positive, negative or neutral, answer in one word: {review}",
                }],
            ),
        )
        for i, review in enumerate(reviews)
    ]
)
print(batch.id, batch.processing_status)  # msgbatch_..., in_progress

When the batch is created its status is in_progress. You check the status periodically; once it is ended, the results are ready:

import time

while client.messages.batches.retrieve(batch.id).processing_status != "ended":
    time.sleep(60)

for result in client.messages.batches.results(batch.id):
    if result.result.type == "succeeded":
        message = result.result.message
        print(result.custom_id, message.content[0].text)
    else:
        print(result.custom_id, "failed:", result.result.type)

Each result has one of four types: succeeded, errored, canceled or expired. You only need to resend the ones that failed.

Lesser-known details

Where first-time Batches API users usually get stuck:

  • Results come back unordered. The result list may not follow the order you sent. Always match by custom_id, never by position.
  • Validation happens later. An invalid parameter doesn't fail when you create the batch; it comes back as errored after processing ends. Before sending a large batch, test the request shape with a single standard API call.
  • custom_id rules are strict: 1 to 64 characters, only letters, digits, hyphens and underscores.
  • A few parameters aren't supported: stream: true, fast mode (speed) and max_tokens: 0 return a validation error inside a batch.
  • Combine it with prompt caching. If every request shares the same long system prompt or document, cache it; the caching and batch discounts stack. Because batches can run longer than 5 minutes, Anthropic recommends the 1-hour cache duration.
  • Batches are scoped to a workspace and can be canceled. If you sent the wrong job, stop it with cancel.

How we use the Batches API at EngerekTech

The blog you are reading is published in both Turkish and English. The English translation feature in our portal sends its translation requests through the Batches API instead of the standard API. Getting a translation in a few minutes instead of a few seconds changes nothing for us, but we save half the cost on every translation.

Decisions like this look small one by one. But in a business that builds AI into its product, deciding up front which requests go out instantly and which go in bulk directly shapes the monthly bill. If you are planning to integrate AI into your own product, get in touch with us; let's build an architecture that accounts for cost from day one.

Frequently asked questions

Are batch results lower quality than the standard API?

No. The same model runs with the same parameters. The only difference is when the answer arrives.

How long can a batch take at most?

Most batches finish in under an hour. Requests that can't be completed within 24 hours are marked expired, and you are not billed for them.

How long can I keep the results?

Results can be downloaded for 29 days after the batch is created. Save anything you want to keep permanently in your own database.

Which models are supported?

All of Anthropic's active models work with the Batches API, including Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 4.5.

Sources

Source: platform.claude.com

ShareLinkedInXWhatsApp
Need help with this?

If you would like to apply what this post covers to your own project, let’s look at it together.

Write to us
YE

Founder of EngerekTech. Builds web, mobile and enterprise software for businesses with Angular, Spring Boot and Flutter, and made the KPSS Düello and Kelime Kavanozu apps. On the blog he covers AI tools and software development as he uses them in his own projects.