Benchmark

IMO-AnswerBench

IMO-AnswerBench is a benchmark for evaluating mathematical reasoning capabilities on International Mathematical Olympiad (IMO) problems, focusing on…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
20

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3.2DeepSeek0.7832025-12-01evidence
DeepSeek-V4-Flash-0423DeepSeek0.8512026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.8842026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.8982026-04-23evidence
GLM-4.7Z.ai0.822025-12-22evidence
GLM-5.1Z.ai0.8382026-04-07evidence
GLM-5.2Z.ai0.912026-06-16evidence
Hy3Tencent0.92026-07-06evidence
Kimi K2-Thinking-0905Moonshot AI0.7862025-09-05evidence
Kimi K2.5Moonshot AI0.8182026-01-27evidence
Kimi K2.6Moonshot AI0.862026-04-20evidence
LongCat-Flash-Thinking-2601Meituan0.7862026-01-14evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.9232026-06-04evidence
Qwen3.5-397B-A17BQwen0.8092026-02-16evidence
Qwen3.6 PlusQwen0.8382026-04-02evidence
Qwen3.6-27BQwen0.8082026-04-21evidence
Qwen3.6-35B-A3BQwen0.7892026-04-16evidence
Qwen3.7 MaxQwen0.92026-05-19evidence
Qwen3.7-PlusQwen0.862026-05-31evidence
Step-3.5-FlashStepFun0.8542026-02-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.