Benchmark

MMLU

Saturated at the frontier and heavily contaminated. Useful as a historical baseline, not as a ranking signal.

Released
2020-09-07
Categories
knowledge
Openness
unknown
Source
model_reports
Reported scores
4

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1DeepSeek90.82025-01-22evidence
DeepSeek-V3DeepSeek88.52024-12-27evidence
Gemini 1.5 ProGoogle85.92024-03-08evidence
Mistral Large 2Mistral84.02024-07-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.