Benchmark

AIME 2024

American Invitational Mathematics Examination 2024, consisting of 30 challenging mathematical reasoning problems from AIME I and AIME II competitions…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
53

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3.7 SonnetAnthropic0.82025-02-24evidence
DeepSeek R1 Distill Llama 70BDeepSeek0.8672025-01-20evidence
DeepSeek R1 Distill Llama 8BDeepSeek0.82025-01-20evidence
DeepSeek R1 Distill Qwen 1.5BDeepSeek0.5272025-01-20evidence
DeepSeek R1 Distill Qwen 14BDeepSeek0.82025-01-20evidence
DeepSeek R1 Distill Qwen 32BDeepSeek0.8332025-01-20evidence
DeepSeek R1 Distill Qwen 7BDeepSeek0.8332025-01-20evidence
DeepSeek R1 ZeroDeepSeek0.8672025-01-20evidence
DeepSeek-R1-0528DeepSeek0.9142025-05-28evidence
DeepSeek-V3DeepSeek0.3922024-12-25evidence
DeepSeek-V3 0324DeepSeek0.5942025-03-25evidence
DeepSeek-V3.1DeepSeek0.6632025-01-10evidence
Gemini 2.0 Flash ThinkingGoogle0.7332025-01-21evidence
Gemini 2.5 FlashGoogle0.882025-05-20evidence
Gemini 2.5 ProGoogle0.922025-05-20evidence
GLM-4.5Z.ai0.912025-07-28evidence
GLM-4.5-AirZ.ai0.8942025-07-28evidence
GPT-4.1OpenAI0.4812025-04-14evidence
GPT-4.1 miniOpenAI0.4962025-04-14evidence
GPT-4.1 nanoOpenAI0.2942025-04-14evidence
GPT-4.5OpenAI0.3672025-02-27evidence
GPT-4oOpenAI0.1312024-08-06evidence
Granite 3.3 8B BaseIBM0.8122025-04-16evidence
Granite 3.3 8B InstructIBM0.8122025-04-16evidence
Grok-3xAI0.9332025-02-17evidence
Grok-3 MinixAI0.9582025-02-17evidence
Kimi K2 0905Moonshot AI0.722025-09-05evidence
Kimi K2 InstructMoonshot AI0.6962025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.6962025-09-05evidence
Kimi-k1.5Moonshot AI0.7752025-01-20evidence
LongCat-Flash-LiteMeituan0.72192026-02-05evidence
LongCat-Flash-ThinkingMeituan0.9332025-09-22evidence
Magistral MediumMistral0.7362025-06-10evidence
Magistral Small 2506Mistral0.70682025-06-10evidence
Min istral 3 (3B Reasoning 2512)Mistral0.7752025-12-04evidence
MiniCPM-SALAOpenBMB0.83752026-02-11evidence
MiniMax M1 40KMiniMax0.8332025-06-16evidence
MiniMax M1 80KMiniMax0.862025-06-16evidence
Ministral 3 (14B Reasoning 2512)Mistral0.8982025-12-04evidence
Ministral 3 (8B Reasoning 2512)Mistral0.862025-12-04evidence
o1OpenAI0.7432024-12-17evidence
o1-previewOpenAI0.422024-09-12evidence
o1-proOpenAI0.862024-12-17evidence
o3OpenAI0.9162025-04-16evidence
o3-miniOpenAI0.8732025-01-30evidence
o4-miniOpenAI0.9342025-04-16evidence
Phi 4 ReasoningMicrosoft0.7532025-04-30evidence
Phi 4 Reasoning PlusMicrosoft0.8132025-04-30evidence
Qwen3 235B A22BQwen0.8572025-04-29evidence
Qwen3 30B A3BQwen0.8042025-04-29evidence
Qwen3 32BQwen0.8142025-04-29evidence
QwQ-32BQwen0.7952025-03-05evidence
QwQ-32B-PreviewQwen0.52024-11-28evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.