Comparison
Compare AI models
Pick up to four models and see every rating side by side.
How the comparison works
Every row in the comparison table comes from the same OpenRating measurements used on the leaderboard. The OpenRating Index condenses multi-step reasoning, knowledge, and autonomous agentic benchmarks into a single 0–100 score. Speed is blended output in tokens per second, and time to first token measures responsiveness under load. Task Cost measures end-to-end compute expenses per benchmark run, alongside standard input/output token pricing.
Reading the numbers
Higher OpenRating Index and Reasoning Index scores are better. Higher output speed and lower time to first token mean a snappier model. Task Cost gives you the true empirical workload cost accounting for chain-of-thought tokens and prompt caching economics. When open-weights and proprietary models score closely, self-hosting feasibility and private data governance are often key deciding factors.