Benchmark

BoolQ

BoolQ is a reading comprehension dataset for yes/no questions containing 15,942 naturally occurring examples. Each example consists of a question…

Modality
text
Categories
reasoning, language
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemma 2 27BGoogle0.8482024-06-27evidence
Gemma 2 9BGoogle0.8422024-06-27evidence
Gemma 3n E2BGoogle0.7642025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.7642025-05-20evidence
Gemma 3n E4BGoogle0.8162025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.8162025-05-20evidence
Hermes 3 70BNous Research0.88042024-08-15evidence
Phi 4 MiniMicrosoft0.8122025-02-01evidence
Phi-3.5-mini-instructMicrosoft0.782024-08-23evidence
Phi-3.5-MoE-instructMicrosoft0.8462024-08-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.