EVL
MMLU
Massive Multitask Language Understanding
What this measures
How much a model simply *knows* — broad academic knowledge across 57 subjects, from history and law to physics, medicine, and economics.
Sample
Which body cavity contains the pituitary gland? (A) Abdominal (B) Cranial (C) Pleural (D) Spinal → Answer: B
Leaderboard · 17 models
Accuracy (% of multiple-choice questions answered correctly) — higher is betterBenchmark health
activeDiscriminating but climbing — the frontier is moving up this benchmark.
87.6
top-10 mean
91.8
best
17
models
What improves this score
Datasets & environments shown to raise it.
Research using this benchmark
all →K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMsNITP: Next Implicit Token Prediction for LLM Pre-trainingLength Penalties Make Chain-of-Thought Less MonitorableHolistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement LearningWho Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
Matched by name in title/abstract.