Benchmark

AIME

pass@k, majority vote and Python tool access each shift this by double digits.

Released
2024-02-01
Categories
math
Openness
unknown
Source
model_reports
Reported scores
10

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude Opus 4Anthropic75.52025-05-22evidence
Claude Sonnet 4Anthropic70.52025-05-22evidence
DeepSeek-R1DeepSeek79.82025-01-22evidence
DeepSeek-V3DeepSeek39.22024-12-27evidence
Gemma 4 (31B)Google89.22026-03-11evidence
GLM-5Z.ai92.72026-02-11evidence
GLM-5.1Z.ai95.32026-04-03evidence
GLM-5.2Z.ai99.22026-06-16evidence
Qwen3-235B-A22B (Thinking)Qwen85.72025-05-14evidence
Qwen3.5-397B-A17BQwen91.32026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.