Benchmark
BBH
Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average…
- Modality
- text
- Categories
- math, reasoning, language
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 12
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| DeepSeek-V2.5 | DeepSeek | 0.843 | 2024-05-08 | evidence |
| ERNIE 4.5 | Baidu | 0.304 | 2025-06-25 | evidence |
| Hermes 3 70B | Nous Research | 0.6782 | 2024-08-15 | evidence |
| MiMo-V2.5-Pro | Xiaomi | 0.884 | 2026-04-27 | evidence |
| MiniCPM-SALA | OpenBMB | 0.8155 | 2026-02-11 | evidence |
| Nova Lite | Amazon | 0.824 | 2024-11-20 | evidence |
| Nova Micro | Amazon | 0.795 | 2024-11-20 | evidence |
| Nova Pro | Amazon | 0.869 | 2024-11-20 | evidence |
| Qwen2 72B Instruct | Qwen | 0.824 | 2024-07-23 | evidence |
| Qwen2.5 14B Instruct | Qwen | 0.782 | 2024-09-19 | evidence |
| Qwen2.5 32B Instruct | Qwen | 0.845 | 2024-09-19 | evidence |
| Qwen3 235B A22B | Qwen | 0.8887 | 2025-04-29 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.