Benchmark

BBH

Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average…

Modality
text
Categories
math, reasoning, language
Openness
unknown
Source
llm_stats
Reported scores
12

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V2.5DeepSeek0.8432024-05-08evidence
ERNIE 4.5Baidu0.3042025-06-25evidence
Hermes 3 70BNous Research0.67822024-08-15evidence
MiMo-V2.5-ProXiaomi0.8842026-04-27evidence
MiniCPM-SALAOpenBMB0.81552026-02-11evidence
Nova LiteAmazon0.8242024-11-20evidence
Nova MicroAmazon0.7952024-11-20evidence
Nova ProAmazon0.8692024-11-20evidence
Qwen2 72B InstructQwen0.8242024-07-23evidence
Qwen2.5 14B InstructQwen0.7822024-09-19evidence
Qwen2.5 32B InstructQwen0.8452024-09-19evidence
Qwen3 235B A22BQwen0.88872025-04-29evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.