citiesabc
Three Things in Your LLM Bill That No Pricing Page Shows
14 Sept 2026

A pricing page gives you two numbers and your invoice is a function of five. The three it leaves out are not fine print — each one can move a bill by more than the difference between two adjacent price tiers. OrcaRouter publishes the two rates alongside real production behaviour, which is what makes the gap between a llm api pricing table and an invoice measurable rather than a suspicion.
Rates read 2026-09-09.
1. How much the model decides to write
Output rate is per token. Token count is a model property, and it varies by more than two to one for the same brief.
On EQ-Bench Creative Writing v3 — an independent LLM-judged benchmark, read 2026-09-08 — the same prompts produced 8,548 output tokens from GPT-5.6 Sol and 6,003 from Claude Opus 5. That is a 42% difference in the quantity you are billed for, and it inverts the price comparison:
| Rate /1M out | Tokens | Cost per piece | |
| GPT-5.6 Sol | $20.00 | 8,548 | $0.171 |
| Claude Opus 5 | $25.00 | 6,003 | $0.150 |
The 20% cheaper rate produces a 14% higher bill. Nothing on either pricing page hints at this, because verbosity is not a published spec.
What to do: log actual output tokens per call from the API response, and compute `mean output tokens ÷ 1e6 × rate` per model. That is your real unit cost. Those benchmark lengths come from open-ended creative prompts specifically, so do not carry the numbers over — carry the method.
2. Your read-to-write mix
The two rates are not weighted equally in your bill; the weighting is set by your workload, and the two extremes barely resemble each other.
A read-heavy call — 20,000 tokens of document in, 500 tokens of summary out, on DeepSeek V4 Flash at $0.24 / $0.73:
• input: 20,000 × $0.24/1M = $0.0048
• output: 500 × $0.73/1M = $0.00037
• input is 93% of the call
A write-heavy call on the same model — 500 in, 6,000 out:
• input: $0.00012
• output: $0.0044
• output is 97% of the call
Same model, same rates, and which published number matters flips completely. This is why the output-to-input ratio deserves a column of its own: it runs from 3.0x (DeepSeek, Qwen3.8 Max, Grok 4.6) to 6.0x (Gemini 3.5 Flash, GPT-5.5 Pro), and a 6.0x model gives back on generation whatever its low input price earned on summarisation.
What to do: pull mean input tokens and mean output tokens per call from your request logs, per endpoint. Most teams already have this and have never looked. It tells you which of the two published columns to sort by, and it is usually not the one they were sorting by.

3. Reasoning tokens, where they are billed as output
Models with configurable effort spend tokens thinking before they answer. Where the provider bills those as output tokens, your effort setting is a pricing lever — and often a bigger one than your model choice.
The arithmetic is unforgiving: if a task needs 500 tokens of visible answer and the model spends 4,000 thinking at max effort, you are paying nine times the naive estimate. Drop to low effort and the same call may cost a fraction, at some accuracy cost you have to measure.
This one is genuinely provider-specific — how effort is exposed, whether thinking tokens are separately metered, and whether they are discounted all vary. The general rule that survives: on any model where you can set effort, treat the setting as part of the price, and measure the token counts at each level on your own task before standardising.
What is actually knowable from a pricing page
To be fair to pricing pages: the two rates are real, they are what you are charged, and they set the ceiling on how cheap a workload can get. The complaint is not that they lie — it is that they are two inputs to a five-input function, and the other three are all properties of your traffic rather than the vendor's price list.
Which means the useful comparison is never fully published by anyone. It is:
```
bill = (mean_in × in_rate + mean_out × out_rate + reasoning_out × out_rate) × calls
```
where the three `mean_*` terms come from your own logs.
A fourth variable, for completeness
Caching, where available, changes the input side materially — repeated prefixes can bill at a fraction of the standard input rate. I have deliberately left it out of the arithmetic above because availability and discount structure vary by provider and I have not verified them across this model set. If your workload has a long fixed system prompt, it is the first thing to check, and it can dominate everything in this article.

The one change that makes all three visible
Everything above is measurable from a single addition to your logging, if it is not there already: record prompt and completion token counts on every call, tagged by endpoint and model.
The API returns both in the response. Most clients discard them. With them stored you can answer, in one query each:
• What is my real per-piece cost per model? Mean completion tokens times output rate. This is trap one, and it is the number your invoice agrees with.
• Which column should I be shopping in? Mean prompt tokens divided by mean completion tokens, per endpoint. Above about 5:1 you are an input-price buyer.
• Are reasoning tokens eating me alive? Where your provider reports them separately, compare them against completion tokens on the same calls. A ratio above 1:1 means your effort setting is the dominant cost lever, not your model choice.
That is three of the five inputs to your bill, from one logging change and three queries. The remaining two — the published rates — you already have.
It is worth saying why this so often is not done: token counts feel like plumbing rather than business data, so they get logged at debug level or not at all. By the time somebody asks what a call actually costs, the data to answer it was never retained. Turning it on is cheap; turning it on retroactively is impossible.
The takeaway
Two published rates, five actual inputs. Verbosity varies by 42% between two models one tier apart and inverts their price ordering. Your read-to-write mix decides whether input is 93% or 3% of a call, which decides which column to sort by. And reasoning tokens, where billed as output, can multiply a call ninefold from a setting rather than a model choice. All three are measurable from your own request logs in an afternoon — and until you have, a pricing table is a rate card, not a forecast.






