Benchmark

BIG-Bench Extra Hard

BIG-Bench Extra Hard (BBEH) is a challenging benchmark that replaces each task in BIG-Bench Hard with a novel task that probes similar reasoning…

Modality
text
Categories
reasoning, language, general
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DiffusionGemma 26B-A4BGoogle0.4762026-06-10evidence
Gemini DiffusionGoogle0.152025-05-20evidence
Gemma 3 12BGoogle0.1632025-03-12evidence
Gemma 3 1BGoogle0.0722025-03-12evidence
Gemma 3 27BGoogle0.1932025-03-12evidence
Gemma 3 4BGoogle0.112025-03-12evidence
Gemma 4 12BGoogle0.532026-05-23evidence
Gemma 4 26B-A4BGoogle0.6482026-04-02evidence
Gemma 4 31BGoogle0.7442026-04-02evidence
Gemma 4 E2BGoogle0.2192026-04-02evidence
Gemma 4 E4BGoogle0.3312026-04-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.