The terminal for AI research.
Which model wins, what it costs to run, how the numbers moved, and the research that moved them — every figure with a source. Press ⌘K to search.
● live · scores, prices & speed verified against Artificial Analysis · Jul 27, 2026
The frontier — right now
full chart · frontier table · cost calculator →Capability vs price
gold staircase = nothing cheaper is better| # | Model | Cap | $/1M | tok/s |
|---|---|---|---|---|
| 1 | Kimi K3 | 99.2 | $6 | 33 |
| 2 | Claude Opus 5 | 97.7 | $10 | 44 |
| 3 | Gemini 3.1 Pro Preview | 97.5 | $4.5 | 132 |
| 4 | GPT-5.5 | 97.0 | $11.25 | — |
| 5 | GPT-5.6 Luna (max) | 96.6 | $2.25 | 171 |
| 6 | Claude Sonnet 5 | 96.4 | $4 | 83 |
| 7 | Gemini 3.6 Flash | 95.7 | $3 | 219 |
| 8 | GLM-5.2 (max) | 95.6 | $2.15 | 157 |
| 9 | Grok 4.5 | 95.6 | $3 | 56 |
| 10 | Claude Fable 5 | 94.9 | $20 | 58 |
This week
the full score ledger →Score movers
14 daysNotable papers
all →ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
▲ 295 on HF · 6d ago
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
▲ 226 on HF · 12d ago
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
▲ 201 on HF · 11d ago
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
▲ 197 on HF · 7d ago
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
▲ 169 on HF · 11d ago
The catalog
cross-linked — every entity connectsRecent activity
full feed →PAPS1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and GenerationyesterdayPAPWhen Does Muon Help Agentic Reinforcement Learning?yesterdayMDLQwen3.6-27B-NVFP4yesterdayMDLBonsai-27B-ggufyesterdayLDRAA Long-Context Reasoning Leaderboard1mo agoLDRTerminal-Bench Hard Leaderboard1mo agoEVLAA Long-Context Reasoning1mo agoEVLTerminal-Bench Hard1mo ago