Skip to content

Long-context LLM pricing

Models that accept 200K tokens or more, ranked by the cost of your workload, and the ones whose price per token rises once a request passes 200K.

Your workload

90M input and 18M output tokens a month (30-day month; token counts are per request).

Assumes: 1,000 conversations a day, each sending about 3,000 tokens of history and instructions and receiving about 600 tokens of replies.

106 of 393 models, sorted by your monthly cost — loading the complete list…

Each row: input / output price, then your monthly cost. Token prices per 1M tokens. Tick up to 4 models to compare.

$3.00a month
$3.02a month
$3.53a month
$4.05a month
$5.04a month
$6.00a month
$6.30a month
$7.56a month
$7.81a month
$8.10a month
$8.10a month
$8.41a month
$9.72a month
$10.13a month
$10.53a month
$11.25a month
$11.34a month
$11.34a month
$11.70a month
$11.70a month
$12.15a month
$12.15a month
$12.60a month
$12.60a month
$12.60a month
$12.64a month
$12.94a month
$14.17a month
$14.22a month
$14.40a month

“Your cost” applies the cache-read price only where the source lists one; a model with no price data shows a dash and is never ranked as cheapest.

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

Models that charge more above 200K tokens

For 11 models, the LiteLLM price table records a second, higher list price that applies when a single request exceeds 200,000 input tokens. A workload that crosses that line pays the higher price, which the monthly cost in the table above does not model.

ModelContextUp to 200K (in / out)Above 200K (in / out)
Anthropic Claude Sonnet 4.51M$3.00 / $15.00$6.00 / $22.50
Google Gemini 2.5 Pro1.05M$1.25 / $10.00$2.50 / $15.00
Google Gemini 3.1 Pro Preview1.05M$2.00 / $12.00$4.00 / $18.00
Google Gemini 3.1 Pro Preview Custom Tools1.05M$2.00 / $12.00$4.00 / $18.00
xAI Grok 4.202M$1.25 / $2.50$2.50 / $5.00
xAI Grok 4.20 Multi-Agent2M$1.25 / $2.50$2.50 / $5.00
xAI Grok 4.31M$1.25 / $2.50$2.50 / $5.00
xAI Grok 4.5500K$2.00 / $6.00$4.00 / $12.00
xAI Grok 4.6500K$2.00 / $6.00$4.00 / $12.00
xAI Grok 4.7500K$2.00 / $6.00$4.00 / $12.00
xAI Grok Build 0.1256K$1.00 / $2.00$2.00 / $4.00

Provider list prices per 1M tokens, as recorded by LiteLLM. Models without a recorded tier may still have one: check the provider’s pricing page.

Paying for a long context

A context window is a ceiling, not a price. What you pay for is the number of tokens you actually send, on every request. A model with a window of a million tokens costs the same as any other for a 3,000-token prompt; it becomes expensive when you fill the window, because the input line of the bill grows with every token and is charged again each time you ask a new question about the same material.

Three ways to keep the bill down

First, send less: retrieve the passages that matter instead of the whole corpus. Second, if the same long material is sent repeatedly, rely on the cache-read price, which the caching guide explains; set “From cache” above to the share that repeats and the ranking adjusts. Third, watch the 200K line: on the models in the table above, one oversized request is billed at the higher tier.

Window and output are different limits

Reading a long document and writing a long document are separate abilities. Open any model in the table to see its maximum output next to its context window: several models with very large windows cap the answer at a small fraction of it.

Questions and answers

Do long prompts cost more per token?

On some models, yes. Several providers charge a higher price per token once a single request passes 200,000 input tokens. The table on this page lists every model for which the LiteLLM price table records such a tier, with both prices.

Is a bigger context window always better?

A bigger window lets you send more, but you pay for every token you send on every request. Sending a whole knowledge base each time is usually dearer than retrieving the few passages that matter, unless the provider’s cache-read price makes the repeated part cheap.

What is the difference between context window and maximum output?

The context window is the total a model can handle in one request, input and output together. Maximum output is the cap on what it can write back. A model with a 1M-token window can still be limited to a much shorter answer.

How do I estimate the cost of a long document?

Count about 1,300 tokens per 1,000 English words, then use the calculator with that number as input tokens. If you ask many questions about the same document, set the cached share to the part that repeats.