Ranking tables sort by capability. Price is not in them. Divide one by the other — what does one point of benchmark cost? — and the order inverts.

Cheapest per point

Limited to models scoring above 30. Below that the work they can do differs enough that the comparison loses meaning.

ModelScoreInputPer point
deepseek-v4-flash-073134.5$0.060$0.0017
glm-5.3-flash41.9$0.090$0.0021
deepseek-v4.1-flash39.5$0.150$0.0038
gpt-5.6-luna37.5$0.200$0.0053
qwen3.8-27b33.9$0.214$0.0063
gemini-3.8-flash41.2$0.750$0.0182
glm-5.344.9$1.400$0.0312

Against the top of the table

ModelScoreInputPer point
claude-fable-5.153.4$10.00$0.1873
gpt-6-astra52.8$10.00$0.1894
claude-opus-550.7$5.00$0.0986

The leader, claude-fable-5.1, costs $0.187 per point. deepseek-v4-flash-0731 costs $0.0017 — 110 times less. In exchange the scores differ by 19 points, 53.4 against 34.5.

What this number can and cannot settle

Cost per point only answers "if both can do the job, which is cheaper." It cannot tell you whether both can do the job. The score is a composite of ten evaluations, so the 19-point gap between 34 and 53 could sit anywhere. If it is in code generation, the cheap model is not an option for code work. If it is in knowledge questions, it may not matter at all for summarisation.

So use this table to narrow candidates. Once two models a hundred times apart in cost per point are on your list, the next step is the per-subject scores — how far apart are they in the subjects your work actually touches. Each model page carries that breakdown.

See model rankingsSee the price table