MMODELYST
Evals/GPQA Diamond
EVL

GPQA Diamond

Graduate-level QA

Source
benchmarkReasoningAccuracy on the hardest 'Diamond' subset198 items
What this measures

Genuinely hard science questions written by PhDs — biology, physics, chemistry — that are 'Google-proof' (you can't just look up the answer). Tests deep reasoning, not recall.

Sample
A graduate-level question like: 'Which quantum number determines the shape of an atomic orbital?' but at a difficulty where even domain experts score ~65%.
Leaderboard · 99 models
Accuracy on the hardest 'Diamond' subset — higher is better
1Gemini 3.1 Pro Preview
94.1
2GPT-5.6 Sol (max)
94.1
3GPT-5.5
93.5
4Kimi K3
93.5
5Claude Opus 5
93.2
6Grok 4.5
93.1
7GPT-5.6 Sol
93.1
8MiniMax-M3
92.9
9Gemini 3.6 Flash
92.8
10Claude Fable 5
92.6
11GPT-5.6 Terra (max)
92.5
12Qwen3.7 Max
92.3
13Gemini 3.5 Flash
92.2
14GPT-5.4
92
15Claude Opus 4.8
92
16GPT-5.3 Codex
91.5
17Claude Opus 4.7
91.4
18Grok 4.20 0309 v2
91.1
19Kimi K2.6
91.1
20Claude Sonnet 5
91.1
21GPT-5.6 Luna (max)
91.1
22Gemini 3 Pro Preview
90.8
23GPT-5.6 Terra
90.8
24GPT-5.2
90.3
25Grok 4.3
90.1
26Qwen3.7 Plus
90
27GPT-5.2 Codex
89.9
28Gemini 3 Flash Preview
89.8
29Muse Spark 1.1
89.8
30Hy3
89.7
31Claude Opus 4.6
89.6
32Kimi K2.7 Code
89.6
33Grok Build 0.1 0616
89.5
34GLM-5.2 (max)
89.5
35GPT-5.6 Luna
89.5
36DeepSeek V4 Flash
89.4
37Qwen3.5 397B A17B
89.3
38Nex-N2-Pro
89.2
39Qwen3.6 Max Preview
88.8
40DeepSeek V4 Pro
88.8
41Grok 4.20 0309
88.5
42Muse Spark
88.4
43Qwen3.6 Plus
88.2
44Kimi K2.5
87.9
45Grok 4
87.7
46Agnes 2.5 Pro Alpha
87.6
47GPT-5.4 mini
87.5
48Claude Sonnet 4.6
87.5
49MiniMax-M2.7
87.4
50GPT-5.1
87.3
51Inkling
87.2
52DeepSeek V3.2 Speciale
87.1
53MiMo-V2-Pro
87
54Motif 3 (Beta)
86.9
55GLM-5.1
86.8
56Hy3-preview
86.7
57Nemotron 3 Ultra 550B A55B
86.7
58Claude Opus 4.5
86.6
59MiMo-V2.5-Pro
86.6
60Qwen3 Max Thinking
86.1
61GPT-5.1 Codex
86
62GLM-4.7
85.9
63Qwen3.5 27B
85.8
64Ring-2.6-1T
85.7
65Qwen3.5 122B A10B
85.7
66Gemma 4 31B
85.7
67KAT Coder Pro V2
85.5
68MiMo-V2-Omni-0327
85.5
69GPT-5
85.4
70Grok 4.1 Fast
85.3
71Nanbeige4.1-3B
84.9
72MiMo-V2.5
84.9
73MiniMax-M2.5
84.8
74GLM-5-Turbo
84.7
75Grok 4 Fast
84.7
76GPT-5.5 Instant (May 2026)
84.6
77MiMo-V2-Flash
84.6
78o3-pro
84.5
79Qwen3.5 35B A3B
84.5
80JT-4.1 Flash 236B A21B
84.5
81Gemini 2.5 Pro
84.4
82Qwen3.6 27B
84.2
83Qwen3.6 35B A3B
84.1
84DeepSeek V3.2
84
85Kimi K2 Thinking
83.8
86Gemini 3.5 Flash-Lite
83.8
87GPT-5 Codex
83.7
88Gemini 2.5 Pro Preview (Mar' 25)
83.6
89MiMo-V2-Flash (Feb 2026)
83.5
90Claude 4.5 Sonnet
83.4
91Step 3.5 Flash
83.1
92MiniMax-M2.1
83
93JT-35B-Flash
82.9
94MiMo-V2-Omni
82.8
95o3
82.7
96Qwen3.5 Omni Plus
82.6
97Step 3.5 Flash 2603
82.6
98GPT-5.5 Instant (May 2026)
82.3
99Gemini 2.5 Pro Preview (May' 25)
82.2
Benchmark health
saturating

Top models cluster near the ceiling — differences here are getting too small to mean much. Weight newer, harder benchmarks more.

93.3
top-10 mean
94.1
best
99
models
Recent changes
Llama 3.2 Instruct 1B17.7 → 20.8yesterday
Gemini 3.6 Flash+92.8yesterday
Gemini 3.5 Flash-Lite+83.8yesterday
Gemma 4 E4B57.6 → 52.2yesterday
Gemma 4 E2B43.3 → 40.5yesterday
What improves this score

Datasets & environments shown to raise it.

+12.0