Benchmark
BIG-Bench Extra Hard
BIG-Bench Extra Hard (BBEH) is a challenging benchmark that replaces each task in BIG-Bench Hard with a novel task that probes similar reasoning…
- Modality
- text
- Categories
- reasoning, language, general
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 11
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| DiffusionGemma 26B-A4B | 0.476 | 2026-06-10 | evidence | |
| Gemini Diffusion | 0.15 | 2025-05-20 | evidence | |
| Gemma 3 12B | 0.163 | 2025-03-12 | evidence | |
| Gemma 3 1B | 0.072 | 2025-03-12 | evidence | |
| Gemma 3 27B | 0.193 | 2025-03-12 | evidence | |
| Gemma 3 4B | 0.11 | 2025-03-12 | evidence | |
| Gemma 4 12B | 0.53 | 2026-05-23 | evidence | |
| Gemma 4 26B-A4B | 0.648 | 2026-04-02 | evidence | |
| Gemma 4 31B | 0.744 | 2026-04-02 | evidence | |
| Gemma 4 E2B | 0.219 | 2026-04-02 | evidence | |
| Gemma 4 E4B | 0.331 | 2026-04-02 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.