Benchmark

HMMT 2025

Harvard-MIT Mathematics Tournament 2025 - A prestigious student-organized mathematics competition for high school students featuring two tournaments…

Modality
text
Categories
math
Openness
unknown
Source
llm_stats
Reported scores
33

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1-0528DeepSeek0.7942025-05-28evidence
DeepSeek-V3.1DeepSeek0.3352025-01-10evidence
DeepSeek-V3.2DeepSeek0.9022025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.9022025-12-01evidence
DeepSeek-V3.2-ExpDeepSeek0.8362025-09-29evidence
DeepSeek-V3.2-SpecialeDeepSeek0.9922025-12-01evidence
GLM-5.1Z.ai0.942026-04-07evidence
GLM-5.2Z.ai0.9442026-06-16evidence
GPT-4.1OpenAI0.2892025-04-14evidence
GPT-4.1 miniOpenAI0.352025-04-14evidence
GPT-5OpenAI0.9332025-08-07evidence
GPT-5 miniOpenAI0.8782025-08-07evidence
GPT-5 nanoOpenAI0.7562025-08-07evidence
GPT-5.2OpenAI0.9942025-12-11evidence
GPT-5.2 ProOpenAI1.02025-12-11evidence
Grok 4 FastxAI0.9332025-08-28evidence
Kimi K2 InstructMoonshot AI0.3882025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.3882025-09-05evidence
Kimi K2-Thinking-0905Moonshot AI0.9752025-09-05evidence
Kimi K2.5Moonshot AI0.9542026-01-27evidence
MiMo-V2-FlashXiaomi0.8442025-12-16evidence
Nemotron 3 Super (120B A12B)NVIDIA0.94732026-03-11evidence
Qwen3.5-122B-A10BQwen0.9142026-02-24evidence
Qwen3.5-27BQwen0.922026-02-24evidence
Qwen3.5-35B-A3BQwen0.892026-02-24evidence
Qwen3.5-397B-A17BQwen0.9482026-02-16evidence
Qwen3.5-4BQwen0.742026-03-02evidence
Qwen3.5-9BQwen0.8322026-03-02evidence
Qwen3.6 PlusQwen0.9672026-04-02evidence
Qwen3.6-27BQwen0.9382026-04-21evidence
Qwen3.6-35B-A3BQwen0.9072026-04-16evidence
Sarvam-105BSarvam AI0.8582026-03-06evidence
Sarvam-30BSarvam AI0.7332026-03-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.