Benchmark

PIQA

PIQA (Physical Interaction: Question Answering) is a benchmark dataset for physical commonsense reasoning in natural language. It tests AI systems'…

Modality
text
Categories
physics, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
ERNIE 4.5Baidu0.5522025-06-25evidence
Gemma 2 27BGoogle0.8322024-06-27evidence
Gemma 2 9BGoogle0.8172024-06-27evidence
Gemma 3n E2BGoogle0.7892025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.7892025-05-20evidence
Gemma 3n E4BGoogle0.812025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.812025-05-20evidence
Hermes 3 70BNous Research0.84442024-08-15evidence
Phi 4 MiniMicrosoft0.7762025-02-01evidence
Phi-3.5-mini-instructMicrosoft0.812024-08-23evidence
Phi-3.5-MoE-instructMicrosoft0.8862024-08-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.