Skip to content

Prompt caching and batch pricing

The two discounts that change an LLM bill the most, with the prices each provider actually publishes and a calculator for your own traffic.

Your workload

180M input and 24M output tokens a month, of which 126M input tokens are read from cache (30-day month; token counts are per request).

Assumes: 5 developers, 200 requests each per day, 6,000 tokens of code context in and 800 out, with 70% of the input re-read from cache.

Only models that publish a cache-read price or a batch variant are listed.

Standard price
$30.00
Every input token at $0.10.
With your cached share
$18.66
70% of input at $0.01 (38% less).
As a batch job
$15.00
Batch variant at $0.05 in and $0.25 out (50% less), no caching applied.

Per month, for OpenAI GPT-6 Luna Pro; token prices per 1M tokens. Each figure uses a price the source publishes: the two discounts are shown separately and are not assumed to stack. Cache-write charges are not included.

Unit
Currency

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

How prompt caching works

Most applications send the same opening to the model again and again: a system prompt, tool definitions, a style guide, a long document the user keeps asking about. With prompt caching the provider keeps the processed form of that opening for a while. The next request that starts with exactly the same tokens re-reads it at a lower price instead of paying full input again.

Three things follow. The saving applies only to the repeated prefix, so put what never changes first and what changes last. The stored prefix expires if it is not reused within the provider’s time window, so caching helps steady traffic far more than occasional calls. And some providers charge a premium to write the prefix the first time, which is recovered only after a few reads.

In the latest collection, 236 models publish a cache-read price. The table below summarises, by provider, how far below the normal input price it sits.

First requestRepeated opening: full input priceNew partNext requestSame opening: cache-read priceNew part
Cached tokens are re-read at the lower price instead of being billed as new input.
ProviderModels with a cache priceMedian discount
OpenAI49 of 6590%
Qwen22 of 5380%
Google20 of 2690%
Anthropic15 of 1590%
Z.ai15 of 1681%
Mistral14 of 1990%
DeepSeek9 of 1490%
xAI7 of 784%
Meta6 of 688%
AionLabs5 of 675%
MiniMax5 of 890%
OpenAI5 of 590%
Xiaomi5 of 599%
Anthropic4 of 493%

Discount of the cache-read price against the input price of the same model on OpenRouter, over paid models. Providers with no published cache-read price are not shown.

How batch pricing works

A batch API takes a file of requests, runs them when capacity is available and returns the answers later, typically within a day. In exchange for giving up the immediate answer, the price is lower. It fits anything that nobody is waiting for: nightly classification, evaluation suites, re-indexing a document base, bulk translation.

OpenRouter lists batch access as a separate variant of a model, and this site folds it into the base model as its batch price. 73 models have one in the latest collection.

ProviderModels with a batch priceMedian discount
OpenAI36 of 6550%
Anthropic14 of 1550%
Google11 of 2650%
Mistral6 of 1950%
Z.ai2 of 1664%
DeepSeek1 of 1463%
MoonshotAI1 of 716%
xAI1 of 720%

Discount of the batch input price against the standard input price of the same model on OpenRouter.

Which one to use

Use caching when requests arrive in real time and share a long prefix: chat with a fixed system prompt, agents with large tool lists, questions about one document. Use batch when the work can wait. If a job is both repetitive and patient, check the provider’s documentation before assuming both discounts apply at once: the rules differ between providers and this site shows only published prices.

To see what either discount does to a whole comparison, set the cached share in the price table or in the calculator: the ranking is recalculated with each model’s own cache-read price.

Questions and answers

What is prompt caching?

When many requests start with the same long prefix, such as a system prompt, a tool list or a document, the provider can store that prefix and charge less to read it again than to process it from scratch. Only the repeated prefix gets the lower price; the new part of each request is billed as normal input.

What is the difference between cache read and cache write?

Reading is the discounted price you pay each time a stored prefix is reused. Writing is what some providers charge the first time the prefix is stored, often more than normal input. Caching pays off when a prefix is reused enough times to cover the write.

What is a batch API?

A way to submit many requests at once and collect the answers later, usually within a day, in exchange for a lower price. It suits work nobody is waiting for: classification, evaluation runs, bulk summaries.

Can caching and batch discounts be combined?

It depends on the provider, and the rules change. The table on this page shows, model by model, the cache-read price and the batch price as the sources list them; it does not assume the two stack.

Why does a model show no cache price?

Because neither OpenRouter nor the LiteLLM table lists one. The model may still cache internally, but with no published price this site does not assume a discount.