CAP
Reasoning
League table — best at reasoning · 30
mean percentile across this capability's benchmarks| # | Model | League | Coverage | $/1M | tok/s |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 99.8 | 2/5 | $10 | 44 |
| 2 | Claude Fable 5 | 99.8 | 2/5 | $20 | 58 |
| 3 | GPT-5.6 Sol (max) | 99.7 | 2/5 | $11.25 | 74 |
| 4 | Gemini 3.1 Pro Preview | 99.4 | 2/5 | $4.5 | 132 |
| 5 | Claude Opus 4.8 | 99.3 | 2/5 | $10 | 63 |
| 6 | GPT-5.6 Sol | 99.1 | 2/5 | $11.25 | 64 |
| 7 | GPT-5.5 | 99.1 | 2/5 | $11.25 | — |
| 8 | Kimi K3 | 99.0 | 2/5 | $6 | 33 |
| 9 | Muse Spark 1.1 | 98.8 | 2/5 | $2 | 124 |
| 10 | GPT-5.6 Terra (max) | 98.6 | 2/5 | $5.63 | 128 |
| 11 | GPT-5.4 | 98.4 | 2/5 | $5.63 | — |
| 12 | Grok 4.5 | 98.4 | 2/5 | $3 | 56 |
| 13 | Gemini 3.5 Flash | 98.4 | 2/5 | $3.38 | 250 |
| 14 | GPT-5.6 Terra | 97.7 | 2/5 | $5.63 | 120 |
| 15 | GPT-5.3 Codex | 97.7 | 2/5 | $4.81 | 126 |
| 16 | GLM-5.2 (max) | 97.6 | 2/5 | $2.15 | 157 |
| 17 | Claude Opus 4.7 | 97.5 | 2/5 | $10 | — |
| 18 | Gemini 3.6 Flash | 97.5 | 2/5 | $3 | 219 |
| 19 | Claude Sonnet 5 | 97.3 | 2/5 | $4 | 83 |
| 20 | Qwen3.7 Max | 97.2 | 2/5 | $3.75 | 200 |
| 21 | Muse Spark | 97.1 | 2/5 | — | — |
| 22 | MiniMax-M3 | 96.9 | 2/5 | $0.525 | 87 |
| 23 | Gemini 3 Pro Preview | 96.8 | 2/5 | $4.5 | — |
| 24 | GPT-5.6 Luna (max) | 96.7 | 2/5 | $2.25 | 171 |
| 25 | Kimi K2.6 | 96.3 | 2/5 | $1.71 | — |
| 26 | Claude Opus 4.6 | 96.2 | 2/5 | $10 | — |
| 27 | Motif 3 (Beta) | 96.2 | 2/5 | — | — |
| 28 | Grok Build 0.1 0616 | 96.0 | 2/5 | $1.25 | — |
| 29 | o3-pro | 95.9 | 1/5 | $35 | — |
| 30 | GPT-5.2 | 95.8 | 2/5 | $4.81 | — |
League = a model's mean percentile across the 5 benchmarksin this capability (best score per benchmark; partial coverage shown). Price & speed via Artificial Analysis.
Benchmarks in this capability (5)
ARC-AGI0 scores
BIG-Bench Hard1480 scores
GPQA Diamond1919 scores
Humanity's Last Exam414 scores
MuSR1480 scores