Prompt caching and batch pricing
The two discounts that change an LLM bill the most, with the prices each provider actually publishes and a calculator for your own traffic.
- Standard price
- $30.00
- Every input token at $0.10.
- With your cached share
- $18.66
- 70% of input at $0.01 (38% less).
- As a batch job
- $15.00
- Batch variant at $0.05 in and $0.25 out (50% less), no caching applied.
Per month, for OpenAI GPT-6 Luna Pro; token prices per 1M tokens. Each figure uses a price the source publishes: the two discounts are shown separately and are not assumed to stack. Cache-write charges are not included.
Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.
How prompt caching works
Most applications send the same opening to the model again and again: a system prompt, tool definitions, a style guide, a long document the user keeps asking about. With prompt caching the provider keeps the processed form of that opening for a while. The next request that starts with exactly the same tokens re-reads it at a lower price instead of paying full input again.
Three things follow. The saving applies only to the repeated prefix, so put what never changes first and what changes last. The stored prefix expires if it is not reused within the provider’s time window, so caching helps steady traffic far more than occasional calls. And some providers charge a premium to write the prefix the first time, which is recovered only after a few reads.
In the latest collection, 236 models publish a cache-read price. The table below summarises, by provider, how far below the normal input price it sits.
| Provider | Models with a cache price | Median discount |
|---|---|---|
| OpenAI | 49 of 65 | 90% |
| Qwen | 22 of 53 | 80% |
| 20 of 26 | 90% | |
| Anthropic | 15 of 15 | 90% |
| Z.ai | 15 of 16 | 81% |
| Mistral | 14 of 19 | 90% |
| DeepSeek | 9 of 14 | 90% |
| xAI | 7 of 7 | 84% |
| Meta | 6 of 6 | 88% |
| AionLabs | 5 of 6 | 75% |
| MiniMax | 5 of 8 | 90% |
| OpenAI | 5 of 5 | 90% |
| Xiaomi | 5 of 5 | 99% |
| Anthropic | 4 of 4 | 93% |
Discount of the cache-read price against the input price of the same model on OpenRouter, over paid models. Providers with no published cache-read price are not shown.
How batch pricing works
A batch API takes a file of requests, runs them when capacity is available and returns the answers later, typically within a day. In exchange for giving up the immediate answer, the price is lower. It fits anything that nobody is waiting for: nightly classification, evaluation suites, re-indexing a document base, bulk translation.
OpenRouter lists batch access as a separate variant of a model, and this site folds it into the base model as its batch price. 73 models have one in the latest collection.
| Provider | Models with a batch price | Median discount |
|---|---|---|
| OpenAI | 36 of 65 | 50% |
| Anthropic | 14 of 15 | 50% |
| 11 of 26 | 50% | |
| Mistral | 6 of 19 | 50% |
| Z.ai | 2 of 16 | 64% |
| DeepSeek | 1 of 14 | 63% |
| MoonshotAI | 1 of 7 | 16% |
| xAI | 1 of 7 | 20% |
Discount of the batch input price against the standard input price of the same model on OpenRouter.
Which one to use
Use caching when requests arrive in real time and share a long prefix: chat with a fixed system prompt, agents with large tool lists, questions about one document. Use batch when the work can wait. If a job is both repetitive and patient, check the provider’s documentation before assuming both discounts apply at once: the rules differ between providers and this site shows only published prices.
To see what either discount does to a whole comparison, set the cached share in the price table or in the calculator: the ranking is recalculated with each model’s own cache-read price.
Questions and answers
What is prompt caching?
When many requests start with the same long prefix, such as a system prompt, a tool list or a document, the provider can store that prefix and charge less to read it again than to process it from scratch. Only the repeated prefix gets the lower price; the new part of each request is billed as normal input.
What is the difference between cache read and cache write?
Reading is the discounted price you pay each time a stored prefix is reused. Writing is what some providers charge the first time the prefix is stored, often more than normal input. Caching pays off when a prefix is reused enough times to cover the write.
What is a batch API?
A way to submit many requests at once and collect the answers later, usually within a day, in exchange for a lower price. It suits work nobody is waiting for: classification, evaluation runs, bulk summaries.
Can caching and batch discounts be combined?
It depends on the provider, and the rules change. The table on this page shows, model by model, the cache-read price and the batch price as the sources list them; it does not assume the two stack.
Why does a model show no cache price?
Because neither OpenRouter nor the LiteLLM table lists one. The model may still cache internally, but with no published price this site does not assume a discount.
Keep going
- Price tableEvery model and its price, sorted by your monthly bill.
- Cost calculatorYour monthly bill on up to four models, line by line.
- Cheapest modelsThe five lowest-cost models in each category.
- Free modelsModels listed at zero price, kept apart from the paid ones.
- Price changesA dated log of every price that went up or down.
- Long-context pricingModels that read very long prompts, and what they charge.