Benchmark

MATH (CoT)

MATH dataset contains 12,500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Llama 3.1 70B InstructMeta0.682024-07-23evidence
Llama 3.1 8B InstructMeta0.5192024-07-23evidence
Ministral 3 (14B Base 2512)Mistral0.6762025-12-04evidence
Ministral 3 (3B Base 2512)Mistral0.6012025-12-04evidence
Ministral 3 (8B Base 2512)Mistral0.6262025-12-04evidence
Mistral Large 3Mistral0.6762025-09-01evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.