Benchmark

AGIEval

A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission…

Modality
text
Categories
legal, math, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
ERNIE 4.5Baidu0.2852025-06-25evidence
Gemma 2 27BGoogle0.5512024-06-27evidence
Gemma 2 9BGoogle0.5282024-06-27evidence
Granite 3.3 8B BaseIBM0.4932025-04-16evidence
Hermes 3 70BNous Research0.56182024-08-15evidence
Ministral 3 (14B Base 2512)Mistral0.6482025-12-04evidence
Ministral 3 (3B Base 2512)Mistral0.5112025-12-04evidence
Ministral 3 (8B Base 2512)Mistral0.5912025-12-04evidence
Ministral 8B InstructMistral0.4832024-10-16evidence
Mistral Small 3 24B BaseMistral0.6582025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.