OpenRating is an independent, open-source AI rating platform. Join the community

OOpenRating

Coding workloads

Best LLM for coding

Ranked on the axes OpenRating measures — reasoning capability, cost per task, and output throughput — across all 25 rated models. Claude Opus 5 (max) leads on reasoning at 89.4.

Read this first

OpenRating does not yet run a coding-specific benchmark. Every figure below comes from the same standardised workload we apply to all models — reasoning, knowledge, and agentic tasks — so this ranking reflects measured reasoning capability and measured economics, not a coding score. We publish it this way rather than invent a metric we have not earned.

The ten strongest reasoning models

Reasoning capability is the closest thing we measure to coding ability. Cost and speed decide whether it is usable at scale.

#ModelProviderReasoningCost / taskSpeedContextLicense
1Claude Opus 5 (max)Anthropic89.4$2.3459 tok/s200KProprietary
2Claude Fable 5 (with fallback)Anthropic87.8$3.1472 tok/s200KProprietary
3Kimi K3 (max)Moonshot AI85.5$0.8438 tok/s256KOpen weights
4GPT-5.6 Sol (max)OpenAI85.0$1.2370 tok/s500KProprietary
5Grok 4.6 (high)xAI83.6$0.8469 tok/s256KProprietary
6GLM-5.3 (max)Zhipu AI81.8$0.6995 tok/s256KOpen weights
7Qwen3.8 2.4T A95BAlibaba79.8$1.0965 tok/s256KOpen weights
8Muse Spark 1.2 (xhigh)Muse AI78.2$0.4098 tok/s128KProprietary
9GPT-5.6 Terra (max)OpenAI76.8$0.5186 tok/s400KProprietary
10Gemini 3.7 Flash (high)Google75.8$0.40392 tok/s1MProprietary

Why cost per task beats price per token

Coding is the workload where token pricing misleads most. An agent working through a repository does not send one prompt and stop — it reads files, writes patches, runs tests, reads the failures, and tries again. Each loop resends context, so a model that looks cheap per million tokens can be the expensive one by the end of a task.

Cost per task measures the whole workload instead: reasoning tokens, answer generation, cache writes, cache hits, and input payload, summed for one standardised task. That is the number that scales with your agent's retry behaviour.

Reasoning tokens are usually the largest line. A model that thinks longer is not being wasteful if it needs fewer attempts.

Estimate your own token costs →

Best reasoning per dollar

Reasoning index divided by task cost. A derived ratio from two measured values — not a benchmark.

GPT-5.6 Luna (max)

OpenAI · $0.048 per task

1475.0

gpt-oss-120b (high)

OpenAI · $0.073 per task

846.6

Muse Glimmer (high)

Muse AI · $0.073 per task

682.2

Nemotron 3.5 Lightning

NVIDIA · $0.076 per task

677.6

Gemini 3.5 Flash-Lite

Google · $0.096 per task

567.7

If your code cannot leave the building

For many teams the deciding constraint is not capability or cost — it is whether proprietary source can be sent to a hosted API at all. Open-weights models can be self-hosted, which removes that question entirely and moves the cost from per-task billing to fixed infrastructure. The trade is real: you take on serving, quantisation, and upgrade work that a hosted API absorbs for you.

How this ranking is built

Every model is scored on the same standardised workload. Reasoning index is a 0–100 composite of multi-step reasoning tasks. Cost per task is the full end-to-end spend for one benchmark task, including reasoning tokens, cache writes and hits, answer generation, and input payload. Speed is blended output generation in tokens per second under standardised load. Nothing on this page is a vendor-reported figure.