Benchmark

MMLU (CoT)

Chain-of-Thought variant of the Massive Multitask Language Understanding benchmark, evaluating language models across 57 tasks including elementary…

Modality
text
Categories
legal, math, reasoning, language, finance, general, healthcare
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Llama 3.1 405B InstructMeta0.8862024-07-23evidence
Llama 3.1 70B InstructMeta0.862024-07-23evidence
Llama 3.1 8B InstructMeta0.732024-07-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.