MMODELYST
CAP

Code

League table — best at code · 30
mean percentile across this capability's benchmarks
#ModelLeagueCoverage$/1Mtok/s
1Claude Fable 5100.01/4$2058
2Gemini 3.1 Pro Preview99.81/4$4.5132
3Kimi K399.51/4$633
4Gemini 3 Pro Preview99.42/4$4.5
5Muse Spark 1.199.31/4$2124
6GPT-5.499.01/4$5.63
7GPT-5.598.51/4$11.25
8GPT-5.6 Sol (max)98.31/4$11.2574
9GPT-5.6 Sol98.11/4$11.2564
10Claude Opus 597.81/4$1044
11GPT-5.2 Codex97.61/4$4.81
12Claude Opus 4.797.31/4$10
13Grok 4.597.11/4$356
14GPT-5.6 Terra (max)96.81/4$5.63128
15Gemini 3 Flash Preview96.82/4$1.13
16GPT-5.296.72/4$4.81
17Claude Sonnet 596.61/4$483
18Kimi K2.696.41/4$1.71
19Claude Opus 4.896.11/4$1063
20GPT-5.3 Codex95.91/4$4.81126
21Gemini 3.5 Flash95.61/4$3.38250
22Gemini 3.6 Flash95.41/4$3219
23GPT-5.6 Luna (max)95.11/4$2.25171
24Claude Opus 4.594.92/4$10
25Claude Opus 4.694.61/4$10
26GPT-5.6 Terra94.41/4$5.63120
27Muse Spark94.21/4
28GLM-5.2 (max)93.71/4$2.15157
29GPT-5.5 Instant (May 2026)93.41/4$11.25
30GLM-4.793.32/4$1

League = a model's mean percentile across the 4 benchmarksin this capability (best score per benchmark; partial coverage shown). Price & speed via Artificial Analysis.

Benchmarks in this capability (4)
HumanEval16 scores
LiveCodeBench279 scores
MBPP0 scores
SciCode412 scores