Coding workloads
Best LLM for coding
Ranked on the axes OpenRating measures — reasoning capability, cost per task, and output throughput — across all 25 rated models. Claude Opus 5 (max) leads on reasoning at 89.4.
Read this first
OpenRating does not yet run a coding-specific benchmark. Every figure below comes from the same standardised workload we apply to all models — reasoning, knowledge, and agentic tasks — so this ranking reflects measured reasoning capability and measured economics, not a coding score. We publish it this way rather than invent a metric we have not earned.
The ten strongest reasoning models
Reasoning capability is the closest thing we measure to coding ability. Cost and speed decide whether it is usable at scale.
| # | Model | Provider | Reasoning | Cost / task | Speed | Context | License |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max) | Anthropic | 89.4 | $2.34 | 59 tok/s | 200K | Proprietary |
| 2 | Claude Fable 5 (with fallback) | Anthropic | 87.8 | $3.14 | 72 tok/s | 200K | Proprietary |
| 3 | Kimi K3 (max) | Moonshot AI | 85.5 | $0.84 | 38 tok/s | 256K | Open weights |
| 4 | GPT-5.6 Sol (max) | OpenAI | 85.0 | $1.23 | 70 tok/s | 500K | Proprietary |
| 5 | Grok 4.6 (high) | xAI | 83.6 | $0.84 | 69 tok/s | 256K | Proprietary |
| 6 | GLM-5.3 (max) | Zhipu AI | 81.8 | $0.69 | 95 tok/s | 256K | Open weights |
| 7 | Qwen3.8 2.4T A95B | Alibaba | 79.8 | $1.09 | 65 tok/s | 256K | Open weights |
| 8 | Muse Spark 1.2 (xhigh) | Muse AI | 78.2 | $0.40 | 98 tok/s | 128K | Proprietary |
| 9 | GPT-5.6 Terra (max) | OpenAI | 76.8 | $0.51 | 86 tok/s | 400K | Proprietary |
| 10 | Gemini 3.7 Flash (high) | 75.8 | $0.40 | 392 tok/s | 1M | Proprietary |
Why cost per task beats price per token
Coding is the workload where token pricing misleads most. An agent working through a repository does not send one prompt and stop — it reads files, writes patches, runs tests, reads the failures, and tries again. Each loop resends context, so a model that looks cheap per million tokens can be the expensive one by the end of a task.
Cost per task measures the whole workload instead: reasoning tokens, answer generation, cache writes, cache hits, and input payload, summed for one standardised task. That is the number that scales with your agent's retry behaviour.
Reasoning tokens are usually the largest line. A model that thinks longer is not being wasteful if it needs fewer attempts.
Estimate your own token costs →Best reasoning per dollar
Reasoning index divided by task cost. A derived ratio from two measured values — not a benchmark.
OpenAI · $0.048 per task
1475.0
OpenAI · $0.073 per task
846.6
Muse AI · $0.073 per task
682.2
NVIDIA · $0.076 per task
677.6
Google · $0.096 per task
567.7
If your code cannot leave the building
For many teams the deciding constraint is not capability or cost — it is whether proprietary source can be sent to a hosted API at all. Open-weights models can be self-hosted, which removes that question entirely and moves the cost from per-task billing to fixed infrastructure. The trade is real: you take on serving, quantisation, and upgrade work that a hosted API absorbs for you.
How this ranking is built
Every model is scored on the same standardised workload. Reasoning index is a 0–100 composite of multi-step reasoning tasks. Cost per task is the full end-to-end spend for one benchmark task, including reasoning tokens, cache writes and hits, answer generation, and input payload. Speed is blended output generation in tokens per second under standardised load. Nothing on this page is a vendor-reported figure.