Benchmark

MATH-500

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
32

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3.7 SonnetAnthropic0.9622025-02-24evidence
DeepSeek R1 Distill Llama 70BDeepSeek0.9452025-01-20evidence
DeepSeek R1 Distill Llama 8BDeepSeek0.8912025-01-20evidence
DeepSeek R1 Distill Qwen 1.5BDeepSeek0.8392025-01-20evidence
DeepSeek R1 Distill Qwen 14BDeepSeek0.9392025-01-20evidence
DeepSeek R1 Distill Qwen 32BDeepSeek0.9432025-01-20evidence
DeepSeek R1 Distill Qwen 7BDeepSeek0.9282025-01-20evidence
DeepSeek R1 ZeroDeepSeek0.9592025-01-20evidence
DeepSeek-V3DeepSeek0.9022024-12-25evidence
DeepSeek-V3 0324DeepSeek0.942025-03-25evidence
GLM-4.5Z.ai0.9822025-07-28evidence
GLM-4.5-AirZ.ai0.9812025-07-28evidence
Granite 3.3 8B BaseIBM0.69022025-04-16evidence
Granite 3.3 8B InstructIBM0.69022025-04-16evidence
Kimi K2 InstructMoonshot AI0.9742025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.9742025-09-05evidence
Kimi-k1.5Moonshot AI0.9622025-01-20evidence
Llama 3.1 Nemotron Nano 8B V1NVIDIA0.9542025-03-18evidence
Llama 3.1 Nemotron Ultra 253B v1NVIDIA0.972025-04-07evidence
Llama-3.3 Nemotron Super 49B v1NVIDIA0.9662025-03-18evidence
LongCat-Flash-ChatMeituan0.9642025-08-29evidence
LongCat-Flash-LiteMeituan0.9682026-02-05evidence
LongCat-Flash-ThinkingMeituan0.9922025-09-22evidence
MiniMax M1 40KMiniMax0.962025-06-16evidence
MiniMax M1 80KMiniMax0.9682025-06-16evidence
Nemotron Nano 9B v2NVIDIA0.9782025-08-18evidence
o1-miniOpenAI0.92024-09-12evidence
Phi 4 Mini ReasoningMicrosoft0.9462025-04-30evidence
QwQ-32BQwen0.9062025-03-05evidence
QwQ-32B-PreviewQwen0.9062024-11-28evidence
Sarvam-105BSarvam AI0.9862026-03-06evidence
Sarvam-30BSarvam AI0.972026-03-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.