When comparing model prices most people look at the input rate. It is the number that appears first in any table. But the bill is often decided by the output rate instead. We measured the gap across 225 chat models.
The median is 4x, the maximum 12x
Dividing the output rate by the input rate gives a median of 4.0x. The lowest is 0.8x — output is actually cheaper — and the highest is 12.2x. These sit side by side in the same table.
| Output vs input | Models |
|---|---|
| 4x | 34 |
| 5x | 33 |
| 6x | 26 |
| 3x | 23 |
| 8x | 19 |
| 2x | 12 |
| 1x | 9 |
116 models — more than half — fall between 3x and 6x. So a rough "about 4x" is not far off. The trouble is the 19 models at 8x. Pick one of those and you pay twice your estimate.
What that means for one job
Assume a job that uses one million input tokens and 200,000 output tokens — roughly the shape of feeding in documents and getting summaries back.
| Model | Input | Output | Ratio | Total |
|---|---|---|---|---|
| claude-fable-5.1 | $10.00 | $50.00 | 5.0x | $20.00 |
| gpt-5.6-sol | $2.00 | $10.00 | 5.0x | $4.00 |
| qwen3.8-max-0902 | $2.00 | $6.00 | 3.0x | $3.20 |
| glm-5.3 | $1.40 | $4.40 | 3.1x | $2.28 |
On input alone claude-fable-5.1 costs seven times glm-5.3. For this job the totals are $20.00 against $2.28 — nine times. The more output-heavy the work, the wider the gap. For jobs that feed in long documents and take short answers, the input rate is the whole story.
What to check
There is no shortcut beyond estimating your own input-to-output ratio and combining both rates with it. Choosing from a single column means paying double your estimate on an 8x model. Our price table shows input and output side by side, and each model page carries a monthly cost estimate.