Pick a machine, or set a budget and let it search 100+ builds for the ones nothing else beats. Every model repositions horizontally by how fast it decodes on the selected build, and drops into the zero column if it will not load. Click any model for its own price-vs-speed staircase on the right. Hardware is filtered to upgrades that actually move quality or speed. Ratings 19 Aug 2026, prices 21 Aug 2026 · sourced in AUD inc GST, switchable to USD.
Every build is compared against everything cheaper, and only the ones that win by a margin worth the money get a card — a quality step big enough to be a different class of model, or a speed step big enough to feel different. The rest are listed underneath with the reason. Build your own in the last slot.
Up and to the right wins. Up = smarter, right = faster. The frontier is what your machine can actually do; everything below-left is dominated by something else you could run instead.
Why models jump when you switch hardware. Decode speed is set almost entirely by which memory tier the
active weights sit in. A GPU with more VRAM does not make arithmetic faster — it moves bytes from the ~85 GB/s
tier to the ~600 GB/s tier. Because 1/(1−f) is hyperbolic, the payoff is back-loaded: covering
50% of experts doubles speed, but covering 90% is a 7.6× gain.
Vertical bars mark models whose score swings with reasoning effort — GLM-5.2 is 34 at default and 51 at max. Plotting one number would hide that.
Dashed rings are modelled speeds; solid dots are published benchmarks. Where a model has been benchmarked on any machine, that correction is carried across to the machines it has not been benchmarked on — otherwise a measured machine and a modelled one are not comparable and an upgrade looks better or worse purely by which one happens to have a benchmark. Scored leave-one-out (each measurement predicted without itself): median error 5%, 90th percentile 29%, worst 36%. Treat modelled values as ±30%.
Params stopped predicting quality. Dense Qwen3.6-27B (37) beats the 397B-A17B flagship (34) at 1/14th the size, and Gemma 4 31B (29) beats Llama 4 Maverick (14) at 1/13th. Never buy capacity for a number of parameters — buy it for a specific model you have benchmarked.
Three lenses, because they disagree. Arena text is style-controlled human preference on ordinary chat — read it as “is it pleasant to talk to”. Arena WebDev is the same method on front-end build tasks, and separates these models over roughly 450 rating points against text arena’s 64. Arena Agent is a standardised score on agentic sessions, where 0 is the field average. They rank these models very differently; none is the truth, so pick the one matching your workload. There is deliberately no blended score — coverage differs per lens (shown beside the selector), and averaging would hide the disagreement worth seeing.
On the Qwen 27B-vs-397B inversion. AA rates Qwen3.6-27B dense at 37 and Qwen3.5-397B-A17B at 34 — but AA also rates Qwen3.5-27B dense at 34, i.e. tied with the 397B in the same generation, so the headline is a generational comparison, not a size one. It is also not an agentic artifact: the 27B wins all ten AA components and the non-agentic gap is larger than the agentic one. Arena disagrees outright — qwen3.5-397b-a17b (1442) beats qwen3.5-27b (1409) by 33 Elo and wins every category. Qwen's own card also puts the 397B ahead on MMLU-Pro, GPQA-D and HLE, which points at a serving/harness discrepancy (AA serves open weights via third-party providers, so quantization is a confound) rather than a weighting one.
Bandwidth does not become tokens. Decode scales roughly as BW0.65, and the gap widens with platform size. The cleanest evidence is one Phoronix run of the same llama.cpp benchmark on two servers: an EPYC 9755 (24× DDR5-6400) does 41.9 t/s on gpt-oss-20b while a Xeon 6980P (24× MRDIMM-8800) does 16.2 — the Xeon has 37% more bandwidth and is 2.6× slower. Same story within one machine: swapping RDIMM for MRDIMM lifted measured bandwidth 32% and decode only 6.5–10.5%. Buy a bandwidth number and you may receive a fifth of it.
| Source | As of | Licence / terms | Republishable |
|---|