Benchmark

MMLU-STEM

STEM-focused subset of the Massive Multitask Language Understanding benchmark, evaluating language models on science, technology, engineering, and…

Modality
text
Categories
math, physics, reasoning, chemistry
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen2.5 14B InstructQwen0.7642024-09-19evidence
Qwen2.5 32B InstructQwen0.8092024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.