Skip to content

Cheapest LLM APIs by category

The five lowest-cost paid models in each category, for the workload you set. Ranked by price alone: this is not a quality ranking.

Your workload

90M input and 18M output tokens a month (30-day month; token counts are per request).

Assumes: 1,000 conversations a day, each sending about 3,000 tokens of history and instructions and receiving about 600 tokens of replies.

Unit
Currency

Any text model

Every paid model that returns text.

#ModelCost / mo
1$2.25
2$3.00
3$3.02
4$3.24
5$3.53

Accepts images

Paid models whose input modalities include images.

#ModelCost / mo
1$3.00
2$4.05
3$5.04
4$6.30
5$7.20

Reasoning models

Paid models that OpenRouter flags as supporting a reasoning mode.

#ModelCost / mo
1$3.00
2$3.02
3$3.24
4$3.53
5$4.05

Context of 200K tokens or more

Paid models with a context window of at least 200,000 tokens.

#ModelCost / mo
1$3.00
2$3.02
3$3.53
4$4.05
5$5.04

Context of 1M tokens or more

Paid models with a context window of at least 1,000,000 tokens.

#ModelCost / mo
1$3.53
2$5.04
3$10.53
4$11.34
5$12.15

Token prices per 1M tokens. The cost applies your workload to every paid model in the category; free variants and entries without price data are left out. The Intelligence Index, where shown, is by Artificial Analysis and relayed by OpenRouter.

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

Why “cheapest” depends on your workload

Two models with the same input price can produce very different bills. A summarisation job sends long documents and gets short answers, so the input price decides almost everything. A writing assistant does the opposite, and the output price dominates. That is why this page ranks by the cost of a workload and not by a single column: switch between the scenarios and watch the lists reorder.

What a low price usually buys

The lowest rows are small models: fast, good at classification, extraction and short answers, weaker at long reasoning and at following intricate instructions. A common pattern is to send the easy majority of requests to a cheap model and only the hard ones to an expensive model, which cuts the bill far more than shaving a few cents from the price of a single model.

Check before you commit

Before moving traffic to the cheapest row, run your own prompts through it and read the answers. Then look at the model’s detail for its context window and maximum output, at the caching and batch prices if your prompts repeat, and at the price changes to see how stable its price has been.

Questions and answers

Is the cheapest model the best value?

No. This page ranks by price only, inside categories defined by a filter: it does not measure the quality of the answers. Where OpenRouter relays the Artificial Analysis Intelligence Index for a model, it is shown so you can weigh price against a third-party quality score.

Why are free models not on this list?

Free variants are rate limited and can be withdrawn without notice, so mixing them with paid models would make the paid ranking useless. They have their own page.

How is “cheapest” decided when input and output prices differ?

By the monthly cost of a workload. The page starts with 1,000 requests a day with 3,000 input and 600 output tokens each; change the numbers and the ranking is recalculated.

What does “reasoning model” mean here?

A model that OpenRouter flags as supporting a reasoning mode. These models can spend extra tokens thinking before they answer, and those tokens are usually billed as output, so their real cost per request is higher than the output length suggests.