GPT-6.1 Sol ProOpenAI
- Input not in cache
- $180
- Output
- $180
- Total per month
- $360
- Per day
- $12.00
- Per request
- $0.012
Describe your traffic once and read the monthly bill on up to four models, line by line, with cached input billed at each model’s published price.
Lowest of your selection: Google Gemini 3.8 Flash, at $135 a month.
Every figure comes from the price each source publishes for the model: nothing is assumed where a price is missing. A 30-day month is used throughout.
Collected from the OpenRouter models API (466 entries). 99 models are cross-checked against provider list prices in the LiteLLM price table.
An API bill for a language model has three lines. Input is everything you send: instructions, conversation history, retrieved documents. Output is what the model writes, at a price that is usually several times higher. Cached input is the part of your input the provider has already seen and can re-read at a discount.
Monthly requests are requests per day times 30. Input tokens per month are monthly requests times input tokens per request; the cached share of them is billed at the cache-read price, and the rest at the input price. Output tokens per month are monthly requests times output tokens per request, billed at the output price. Each product is divided by one million, because prices are quoted per million tokens.
If you already call an API, read the token counts from the usage field of a few real responses and average them: guesses are usually off by a wide margin. If you are planning, start from one of the scenarios and adjust. For chat, remember that every turn re-sends the conversation so far, so input grows with the length of the dialogue.
The total is an estimate of the usage charge only. It leaves out cache-write fees, tool and search fees, images and audio, taxes and any committed-use discount. Use it to compare models on equal terms and to see which of the three lines dominates your bill, then confirm on the provider’s pricing page.
Requests per day times 30 gives the requests in a month. Those are multiplied by the input and output tokens per request, divided by one million and multiplied by the model’s price per million tokens. The share of input you mark as cached is billed at the model’s cache-read price when the source lists one.
For English text, a token is about four characters, so 1,000 tokens are roughly 750 words and a full page of prose is 600 to 700 tokens. Code, other languages and long numbers use more tokens per word. Tokenisers differ between providers, so treat any estimate as a range.
Cache-write charges, per-request fees, reasoning tokens that some models bill as output, image and audio inputs, web-search tool fees, taxes and volume discounts. For reasoning models, raise the output tokens to include the hidden reasoning your requests usually produce.
Because the source lists no cache-read price for them. The calculator then bills cached input at the normal input price instead of assuming a discount the provider may not offer.
Yes. The numbers you enter and the models you pick are written into the address of the page. Copy the link and whoever opens it sees the same calculation with the latest collected prices.