Cheapest LLM APIs by category
The five lowest-cost paid models in each category, for the workload you set. Ranked by price alone: this is not a quality ranking.
Any text model
Every paid model that returns text.
| # | Model | Cost / mo |
|---|---|---|
| 1 | $2.25 | |
| 2 | $3.00 | |
| 3 | $3.02 | |
| 4 | $3.24 | |
| 5 | $3.53 |
Accepts images
Paid models whose input modalities include images.
| # | Model | Cost / mo |
|---|---|---|
| 1 | $3.00 | |
| 2 | $4.05 | |
| 3 | $5.04 | |
| 4 | $6.30 | |
| 5 | $7.20 |
Reasoning models
Paid models that OpenRouter flags as supporting a reasoning mode.
| # | Model | Cost / mo |
|---|---|---|
| 1 | $3.00 | |
| 2 | $3.02 | |
| 3 | $3.24 | |
| 4 | $3.53 | |
| 5 | $4.05 |
Context of 200K tokens or more
Paid models with a context window of at least 200,000 tokens.
| # | Model | Cost / mo |
|---|---|---|
| 1 | $3.00 | |
| 2 | $3.02 | |
| 3 | $3.53 | |
| 4 | $4.05 | |
| 5 | $5.04 |
Context of 1M tokens or more
Paid models with a context window of at least 1,000,000 tokens.
| # | Model | Cost / mo |
|---|---|---|
| 1 | $3.53 | |
| 2 | $5.04 | |
| 3 | $10.53 | |
| 4 | $11.34 | |
| 5 | $12.15 |
Token prices per 1M tokens. The cost applies your workload to every paid model in the category; free variants and entries without price data are left out. The Intelligence Index, where shown, is by Artificial Analysis and relayed by OpenRouter.
Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.
Why “cheapest” depends on your workload
Two models with the same input price can produce very different bills. A summarisation job sends long documents and gets short answers, so the input price decides almost everything. A writing assistant does the opposite, and the output price dominates. That is why this page ranks by the cost of a workload and not by a single column: switch between the scenarios and watch the lists reorder.
What a low price usually buys
The lowest rows are small models: fast, good at classification, extraction and short answers, weaker at long reasoning and at following intricate instructions. A common pattern is to send the easy majority of requests to a cheap model and only the hard ones to an expensive model, which cuts the bill far more than shaving a few cents from the price of a single model.
Check before you commit
Before moving traffic to the cheapest row, run your own prompts through it and read the answers. Then look at the model’s detail for its context window and maximum output, at the caching and batch prices if your prompts repeat, and at the price changes to see how stable its price has been.
Questions and answers
Is the cheapest model the best value?
No. This page ranks by price only, inside categories defined by a filter: it does not measure the quality of the answers. Where OpenRouter relays the Artificial Analysis Intelligence Index for a model, it is shown so you can weigh price against a third-party quality score.
Why are free models not on this list?
Free variants are rate limited and can be withdrawn without notice, so mixing them with paid models would make the paid ranking useless. They have their own page.
How is “cheapest” decided when input and output prices differ?
By the monthly cost of a workload. The page starts with 1,000 requests a day with 3,000 input and 600 output tokens each; change the numbers and the ranking is recalculated.
What does “reasoning model” mean here?
A model that OpenRouter flags as supporting a reasoning mode. These models can spend extra tokens thinking before they answer, and those tokens are usually billed as output, so their real cost per request is higher than the output length suggests.
Keep going
- Price tableEvery model and its price, sorted by your monthly bill.
- Cost calculatorYour monthly bill on up to four models, line by line.
- Free modelsModels listed at zero price, kept apart from the paid ones.
- Price changesA dated log of every price that went up or down.
- Caching and batch discountsHow the two discounts work and what you save.
- Long-context pricingModels that read very long prompts, and what they charge.