Skip to content

NVIDIA API pricing

11 NVIDIA models, ranked by what your workload would cost per month. Prices as listed on OpenRouter at 3 Oct 2026, 08:40 UTC.

Your workload

90M input and 18M output tokens a month (30-day month; token counts are per request).

Assumes: 1,000 conversations a day, each sending about 3,000 tokens of history and instructions and receiving about 600 tokens of replies.

11 of 11 models, sorted by your monthly cost — loading the complete list…

Each row: input / output price, then your monthly cost. Token prices per 1M tokens. Tick up to 4 models to compare.

Freea month
Freea month
Freea month
Freea month
Freea month
$8.10a month
$8.41a month
$15.30a month
$21.60a month
$84.60a month
—a month

“Your cost” applies the cache-read price only where the source lists one; a model with no price data shows a dash and is never ranked as cheapest.

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

NVIDIA prices at a glance

  • 11 models listed, 5 of them with a paid price. Batch variants are folded into their base model.
  • Lowest input price: Nemotron 3 Nano 30B A3B, at $0.05 per million input tokens and $0.20 per million output tokens.
  • Highest input price: Nemotron 3 Ultra, at $0.50 input and $2.20 output per million tokens.
  • 3 models publish a cache-read price; the median discount on input is 50% (from 40% to 80%).

These are OpenRouter prices. Open a model in the table to see the provider list price recorded by LiteLLM, where there is one, and whether the two differ. To understand the discounts, read the guide to prompt caching and batch pricing.

Questions and answers

How many NVIDIA models are listed?

11 at the collection of 3 Oct 2026, 08:40 UTC, of which 5 have a paid price on OpenRouter. Batch variants are folded into their base model.

What is the cheapest NVIDIA model?

By input price, Nemotron 3 Nano 30B A3B: $0.05 per million input tokens and $0.20 per million output tokens. Whether it is the cheapest for you depends on your mix of input and output, which is what the table on this page calculates.

What is the most expensive NVIDIA model?

By input price, Nemotron 3 Ultra: $0.50 per million input tokens and $2.20 per million output tokens.

Do NVIDIA models have cached-input prices?

3 of the 11 listed models publish a cache-read price. The table shows it in the “Cached in” column.

Are these the prices NVIDIA charges directly?

They are the prices on OpenRouter. Where LiteLLM records the provider list price for a model, opening the model shows both, and the row is flagged when they differ by more than 1%.