Skip to content

Z.ai API pricing

16 Z.ai models, ranked by what your workload would cost per month. Prices as listed on OpenRouter at 3 Oct 2026, 08:40 UTC.

Your workload

90M input and 18M output tokens a month (30-day month; token counts are per request).

Assumes: 1,000 conversations a day, each sending about 3,000 tokens of history and instructions and receiving about 600 tokens of replies.

10 of 10 models, sorted by your monthly cost — loading the complete list…

Each row: input / output price, then your monthly cost. Token prices per 1M tokens. Tick up to 4 models to compare.

$12.64a month
$22.50a month
$27.00a month
$43.20a month
$55.80a month
$70.20a month
$86.40a month
$88.56a month
$93.60a month
$93.60a month

“Your cost” applies the cache-read price only where the source lists one; a model with no price data shows a dash and is never ranked as cheapest.

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

Z.ai prices at a glance

  • 16 models listed, 16 of them with a paid price. Batch variants are folded into their base model.
  • Lowest input price: GLM 4.7 Flash, at $0.06 per million input tokens and $0.40 per million output tokens.
  • Highest input price: GLM 5.3 Prime, at $2.80 input and $8.80 output per million tokens.
  • 15 models publish a cache-read price; the median discount on input is 81% (from 37% to 90%).
  • 2 models have a batch variant on OpenRouter, with a median discount of 64% on input.

These are OpenRouter prices. Open a model in the table to see the provider list price recorded by LiteLLM, where there is one, and whether the two differ. To understand the discounts, read the guide to prompt caching and batch pricing.

Questions and answers

How many Z.ai models are listed?

16 at the collection of 3 Oct 2026, 08:40 UTC, of which 16 have a paid price on OpenRouter. Batch variants are folded into their base model.

What is the cheapest Z.ai model?

By input price, GLM 4.7 Flash: $0.06 per million input tokens and $0.40 per million output tokens. Whether it is the cheapest for you depends on your mix of input and output, which is what the table on this page calculates.

What is the most expensive Z.ai model?

By input price, GLM 5.3 Prime: $2.80 per million input tokens and $8.80 per million output tokens.

Do Z.ai models have cached-input prices?

15 of the 16 listed models publish a cache-read price. The table shows it in the “Cached in” column.

Are these the prices Z.ai charges directly?

They are the prices on OpenRouter. Where LiteLLM records the provider list price for a model, opening the model shows both, and the row is flagged when they differ by more than 1%.