Benchmark

AIME 2026

All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
21

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DiffusionGemma 26B-A4BGoogle0.6912026-06-10evidence
Gemma 4 12BGoogle0.7752026-05-23evidence
Gemma 4 26B-A4BGoogle0.8832026-04-02evidence
Gemma 4 31BGoogle0.8922026-04-02evidence
Gemma 4 E2BGoogle0.3752026-04-02evidence
Gemma 4 E4BGoogle0.4252026-04-02evidence
GLM-5.1Z.ai0.9532026-04-07evidence
GLM-5.2Z.ai0.9922026-06-16evidence
Inkling-SmallThinking Machines Lab0.9552026-07-30evidence
Kimi K2.6Moonshot AI0.9642026-04-20evidence
MAI-Code-1-FlashMicrosoft0.9252026-06-02evidence
MAI-Thinking-1Microsoft0.9452026-06-02evidence
Muse Glimmer-30BMeta0.9472026-08-10evidence
Qwen3.5-397B-A17BQwen0.9132026-02-16evidence
Qwen3.6 PlusQwen0.9532026-04-02evidence
Qwen3.6-27BQwen0.9412026-04-21evidence
Qwen3.6-35B-A3BQwen0.9272026-04-16evidence
Sakana NamazuSakana AI0.96672026-08-03evidence
Seed 2.0 LiteByteDance0.8832026-02-14evidence
Seed 2.0 ProByteDance0.9422026-02-14evidence
Solar Pro 4Upstage0.9532026-08-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.