Ranking tables sort by capability. Price is not in them. Divide one by the other — what does one point of benchmark cost? — and the order inverts.
Cheapest per point
Limited to models scoring above 30. Below that the work they can do differs enough that the comparison loses meaning.
| Model | Score | Input | Per point |
|---|---|---|---|
| deepseek-v4-flash-0731 | 34.5 | $0.060 | $0.0017 |
| glm-5.3-flash | 41.9 | $0.090 | $0.0021 |
| deepseek-v4.1-flash | 39.5 | $0.150 | $0.0038 |
| gpt-5.6-luna | 37.5 | $0.200 | $0.0053 |
| qwen3.8-27b | 33.9 | $0.214 | $0.0063 |
| gemini-3.8-flash | 41.2 | $0.750 | $0.0182 |
| glm-5.3 | 44.9 | $1.400 | $0.0312 |
Against the top of the table
| Model | Score | Input | Per point |
|---|---|---|---|
| claude-fable-5.1 | 53.4 | $10.00 | $0.1873 |
| gpt-6-astra | 52.8 | $10.00 | $0.1894 |
| claude-opus-5 | 50.7 | $5.00 | $0.0986 |
The leader, claude-fable-5.1, costs $0.187 per point. deepseek-v4-flash-0731 costs $0.0017 — 110 times less. In exchange the scores differ by 19 points, 53.4 against 34.5.
What this number can and cannot settle
Cost per point only answers "if both can do the job, which is cheaper." It cannot tell you whether both can do the job. The score is a composite of ten evaluations, so the 19-point gap between 34 and 53 could sit anywhere. If it is in code generation, the cheap model is not an option for code work. If it is in knowledge questions, it may not matter at all for summarisation.
So use this table to narrow candidates. Once two models a hundred times apart in cost per point are on your list, the next step is the per-subject scores — how far apart are they in the subjects your work actually touches. Each model page carries that breakdown.