MMODELYST
Evals/Humanity's Last Exam
EVL

Humanity's Last Exam

benchmarkReasoningaccuracy %
Leaderboard · 100 models
accuracy % — higher is better
1Claude Fable 5
53.3
2Claude Opus 5
52.6
3GPT-5.6 Sol (max)
47.2
4Claude Opus 4.8
45.7
5Muse Spark 1.1
45.1
6Gemini 3.1 Pro Preview
44.7
7GPT-5.6 Sol
44.7
8GPT-5.5
44.3
9Kimi K3
44.3
10GPT-5.6 Terra (max)
41.8
11GPT-5.4
41.6
12Gemini 3.5 Flash
41
13Grok 4.5
40.3
14GLM-5.2 (max)
40.1
15GPT-5.6 Terra
40
16Muse Spark
39.9
17GPT-5.3 Codex
39.9
18Claude Opus 4.7
39.6
19Claude Sonnet 5
39.6
20Gemini 3.6 Flash
38.3
21Motif 3 (Beta)
38.2
22Qwen3.7 Max
38.1
23Gemini 3 Pro Preview
37.2
24GPT-5.6 Luna (max)
37.2
25MiniMax-M3
37.1
26Claude Opus 4.6
36.7
27Grok Build 0.1 0616
36
28Kimi K2.6
35.9
29DeepSeek V4 Pro
35.9
30GPT-5.6 Luna
35.6
31GPT-5.2
35.4
32Grok 4.3
35
33Gemini 3 Flash Preview
34.7
34MiMo-V2.5-Pro
33.8
35GPT-5.2 Codex
33.5
36Qwen3.7 Plus
33.4
37KAT-Coder-Pro V1
33.4
38Kimi K2.7 Code
32.8
39Nex-N2-Pro
32.4
40Grok 4.20 0309 v2
32.2
41LongCat 2.0
32.1
42DeepSeek V4 Flash
32.1
43Agnes 2.5 Pro Alpha
31.9
44Hy3
31.6
45Grok 4.20 0309
30
46Claude Sonnet 4.6
30
47Inkling
29.7
48Kimi K2.5
29.4
49Qwen3.6 Max Preview
28.9
50Claude Opus 4.5
28.4
51MiMo-V2-Pro
28.3
52MiniMax-M2.7
28.1
53GLM-5.1
28
54Qwen3.5 397B A17B
27.3
55GLM-5
27.2
56GPT-5.4 mini
26.6
57Nemotron 3 Ultra 550B A55B
26.6
58GPT-5.1
26.5
59GPT-5
26.5
60GPT-5.4 nano
26.5
61Qwen3 Max Thinking
26.2
62DeepSeek V3.2 Speciale
26.1
63Qwen3.6 Plus
25.7
64GPT-5 Codex
25.6
65Hy3-preview
25.5
66GLM-5-Turbo
25.4
67MiMo-V2.5
25.2
68GLM-4.7
25.1
69Grok 4
23.9
70GPT-5.1 Codex
23.4
71Qwen3.5 122B A10B
23.4
72Gemma 4 31B
22.7
73Step 3.5 Flash 2603
22.6
74Kimi K2 Thinking
22.3
75Qwen3.5 27B
22.2
76MiniMax-M2.1
22.2
77DeepSeek V3.2
22.2
78Qwen3.6 27B
21.6
79MiMo-V2-Flash
21.1
80Gemini 2.5 Pro
21.1
81MiMo-V2-Omni-0327
20.4
82GPT-5.5 Instant (May 2026)
20.3
83Qwen3.6 35B A3B
20.2
84MiMo-V2-Flash (Feb 2026)
20
85o3
20
86MiMo-V2-Omni
19.9
87Step 3.7 Flash
19.9
88Qwen3.5 35B A3B
19.7
89NVIDIA Nemotron 3 Super 120B A12B
19.2
90Step 3.5 Flash
19.1
91MiniMax-M2.5
19.1
92GPT-5.5 Instant (May 2026)
18.6
93gpt-oss-120b
18.5
94Ring-2.6-1T
18.3
95Gemma 4 26B A4B
18.3
96Grok 4.1 Fast
17.6
97o4-mini
17.5
98Gemini 3.5 Flash-Lite
17.5
99Claude 4.5 Sonnet
17.3
100Gemini 2.5 Pro Preview (Mar' 25)
17.1
Benchmark health
headroom

Plenty of room at the top — this benchmark still separates frontier models clearly.

46.4
top-10 mean
53.3
best
100
models
Recent changes
Gemini 3.6 Flash+38.3yesterday
Gemini 3.5 Flash-Lite+17.5yesterday
Claude Opus 5+52.6yesterday
Motif 3 (Beta)+38.2yesterday
G9v3-3B+4yesterday
What improves this score

Datasets & environments shown to raise it.

No data yet.