Benchmark

HellaSwag

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but…

Released
2019-05-19
Modality
text
Categories
reasoning
Openness
restricted
Source
llm_stats
Reported scores
27

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3 HaikuAnthropic0.8592024-03-13evidence
Claude 3 OpusAnthropic0.9542024-02-29evidence
Claude 3 SonnetAnthropic0.892024-02-29evidence
Command R+Cohere0.8862024-08-30evidence
ERNIE 4.5Baidu0.332025-06-25evidence
Gemini 1.5 FlashGoogle0.8652024-05-01evidence
Gemini 1.5 ProGoogle0.9332024-05-01evidence
Gemma 2 27BGoogle0.8642024-06-27evidence
Gemma 2 9BGoogle0.8192024-06-27evidence
Gemma 3n E2BGoogle0.7222025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.7222025-05-20evidence
Gemma 3n E4BGoogle0.7862025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.7862025-05-20evidence
GPT-4OpenAI0.9532023-06-13evidence
Granite 3.3 8B BaseIBM0.8012025-04-16evidence
Hermes 3 70BNous Research0.88192024-08-15evidence
Llama 3.1 Nemotron 70B InstructNVIDIA0.85582024-10-01evidence
Llama 3.2 3B InstructMeta0.6982024-09-25evidence
MiMo-V2.5-ProXiaomi0.8982026-04-27evidence
Mistral NeMo InstructMistral0.8352024-07-18evidence
Phi 4 MiniMicrosoft0.6912025-02-01evidence
Phi-3.5-mini-instructMicrosoft0.6942024-08-23evidence
Phi-3.5-MoE-instructMicrosoft0.8382024-08-23evidence
Qwen2 72B InstructQwen0.8762024-07-23evidence
Qwen2.5 32B InstructQwen0.8522024-09-19evidence
Qwen2.5-Coder 32B InstructQwen0.832024-09-19evidence
Qwen2.5-Coder 7B InstructQwen0.7682024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.