The score ledger
Changes
Every benchmark score on Modelyst is tracked in an append-only ledger. When a number moves between refreshes — a re-evaluation, a silent endpoint update, a harness change upstream — it lands here, with the old value, the new value, and the date. Last refresh: Jul 27, 2026.
Capability · price · speed drift — last 30 days
GPT-5 minicap -5.1-102 tok/sQwen3.5 4Bcap -5.0+6 tok/sQwen3.5 2Bcap -4.2price −0.04-36 tok/sGPT-5.5 Instant (May 2026)cap +2.7LFM2 2.6Bcap +2.4-336 tok/sQwen3.5 0.8Bcap -1.8price −0.02-30 tok/sMiMo-V2-Procap -1.8price −1.5-48 tok/sMercury 2cap -1.8+113 tok/sHy3-previewcap -1.8-41 tok/sDiffusionGemma 26B A4Bcap -1.8Mistral Medium 3.5cap -1.8-32 tok/sQwen3.5 Omni Pluscap -1.8-3 tok/s
Jul 27, 2026
35 changes · 34 first observationsJul 20, 2026
10 changes · 19 first observationsJul 6, 2026
14 changes · 16 first observationsJun 29, 2026
10 changes · 10 first observationsSource of each value: see the score's provenance on its model page. Scores via Artificial Analysis are medians across providers; changes can reflect re-evaluation, endpoint updates, or methodology changes upstream.