Benchmark

MMLU-Base

Base version of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary mathematics, US…

Modality
text
Categories
legal, math, reasoning, language, finance, general, healthcare
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen2.5-Coder 7B InstructQwen0.682024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.