Skip to content

Claude Sonnet 5.5 vs Llama 4 Maverick pricing

Anthropic Claude Sonnet 5.5 and Meta Llama 4 Maverick side by side: token prices, cache and batch prices, limits, and the monthly cost of your own workload.

Your workload

90M input and 18M output tokens a month (30-day month; token counts are per request).

Assumes: 1,000 conversations a day, each sending about 3,000 tokens of history and instructions and receiving about 600 tokens of replies.

For this workload Meta Llama 4 Maverick costs $28.62 a month and Anthropic Claude Sonnet 5.5 costs $360: 13 times as much.

AttributeClaude Sonnet 5.5Llama 4 Maverick
ProviderAnthropicMeta
Input, per 1M tokens$2.00$0.188
Output, per 1M tokens$10.00$0.652
Cached input (read)$0.20not listed
Batch (input / output)$1.00 / $5.00not listed
Provider list price (input / output)$2.00 / $10.00not recorded
Context window1,000,0001,048,576
Maximum output128,00016,384
Acceptstext, image, filetext, image
Intelligence Index (Artificial Analysis)56not listed
Your cost per month$360$28.62lowest
Unit
Currency

The cost uses each model’s published cache-read price for the cached share and normal input price where none is published. Full detail: Claude Sonnet 5.5 pricing and Llama 4 Maverick pricing.

Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.

How to read this comparison

A lower price per token does not settle which model is cheaper for you. The two models can differ in how the bill splits between input and output: one may have cheaper input and dearer output than the other. Which side wins depends on how long your prompts are compared with the answers you ask for, so set the workload above to your own traffic before reading the last row.

If your requests repeat a long prefix, such as a system prompt or a document, raise “From cache”. The share you set is billed at each model’s cache-read price when one is published, which can change the result when only one of the two offers it. Reasoning models bill hidden thinking as output: for those, raise the output tokens to what your requests really consume.

Price is one axis. Context window, maximum output and accepted inputs are listed because they decide whether a model can do the job at all; the Intelligence Index, where shown, is a third-party score from Artificial Analysis and not a judgement by this site.

Related comparisons

Questions and answers

Which is cheaper, Anthropic Claude Sonnet 5.5 or Meta Llama 4 Maverick?

For 1,000 requests a day with 3,000 input and 600 output tokens each, Anthropic Claude Sonnet 5.5 costs $360 a month and Meta Llama 4 Maverick costs $28.62, so Meta Llama 4 Maverick is cheaper for that workload. A different mix of input and output can change the answer; use the simulator on this page.

What are the prices per million tokens?

Anthropic Claude Sonnet 5.5: $2.00 input and $10.00 output. Meta Llama 4 Maverick: $0.188 input and $0.652 output. These are the OpenRouter prices at the collection of 3 Oct 2026, 08:40 UTC.

Which one has the larger context window?

Meta Llama 4 Maverick, with 1,048,576 tokens against 1,000,000.

Do both offer a cached-input price?

Anthropic Claude Sonnet 5.5: $0.20 per million cached input tokens. Meta Llama 4 Maverick: no cache-read price listed.

Does this comparison say which model is better?

No. It compares prices and limits. Quality depends on your task; where OpenRouter relays the Artificial Analysis Intelligence Index for a model, the table shows it as a third-party reference.