MMODELYST
Capabilities/Instruction Following
CAP

Instruction Following

League table — best at instruction following · 30
mean percentile across this capability's benchmarks
#ModelLeagueCoverage$/1Mtok/s
1Grok 4.20 0309100.01/1$3
2MiniMax-M399.71/1$0.52587
3Nemotron 3 Ultra 550B A55B99.41/1$1.18180
4Grok 4.399.11/1$1.56
5Grok 4.20 0309 v298.81/1$3
6Qwen3.7 Max98.51/1$3.75200
7Nemotron Cascade 2 30B A3B98.21/1
8MiMo-V2.5-Pro97.91/1$0.54465
9DeepSeek V4 Flash97.61/1$0.175118
10Nova 2.0 Pro Preview97.31/1$3.44136
11Qwen3.5 397B A17B97.01/1$1.3566
12Gemini 3 Flash Preview96.71/1$1.13
13Qwen3.7 Plus96.41/1$0.753
14GPT-5.2 Codex96.01/1$4.81
15Gemini 3.1 Flash-Lite95.71/1$0.563299
16Gemini 3.1 Pro Preview95.41/1$4.5132
17Qwen3.6 Max Preview95.11/1$2.92
18DeepSeek V4 Pro94.81/1$0.54471
19GLM-5.194.51/1$2.14
20Gemini 3.5 Flash94.21/1$3.38250
21Kimi K2.693.91/1$1.71
22GPT-5.4 nano93.61/1$0.463
23GPT-5.593.31/1$11.25
24Muse Spark93.01/1
25MiniMax-M2.792.71/1$0.525
26Qwen3.5 122B A10B92.41/1$1.1135
27Qwen3.5 27B92.11/1$0.825
28Gemma 4 31B91.81/135
29GPT-5.291.51/1$4.81
30GPT-5.3 Codex91.21/1$4.81126

League = a model's mean percentile across the 1 benchmarkin this capability (best score per benchmark; partial coverage shown). Price & speed via Artificial Analysis.

Benchmarks in this capability (1)
IFBench330 scores