Benchmark

Arena Hard

Arena-Hard-Auto is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder…

Modality
text
Categories
reasoning, general, creativity, writing
Openness
unknown
Source
llm_stats
Reported scores
26

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V2.5DeepSeek0.7622024-05-08evidence
Granite 3.3 8B BaseIBM0.57562025-04-16evidence
Granite 3.3 8B InstructIBM0.57562025-04-16evidence
IBM Granite 4.0 Tiny PreviewIBM0.2672025-05-02evidence
Jamba 1.5 LargeAI21 Labs0.6542024-08-22evidence
Jamba 1.5 MiniAI21 Labs0.4612024-08-22evidence
Llama-3.3 Nemotron Super 49B v1NVIDIA0.8832025-03-18evidence
MiniStral 3 (14B Instruct 2512)Mistral0.5512025-12-04evidence
Ministral 3 (3B Instruct 2512)Mistral0.3052025-12-04evidence
Ministral 3 (8B Instruct 2512)Mistral0.5092025-12-04evidence
Ministral 8B InstructMistral0.7092024-10-16evidence
Mistral Large 3Mistral0.5512025-09-01evidence
Mistral Small 3 24B InstructMistral0.8762025-01-30evidence
Mistral Small 3.2 24B InstructMistral0.4312025-06-20evidence
Mistral Small 4Mistral0.5832026-03-16evidence
Phi 4Microsoft0.7542024-12-12evidence
Phi 4 MiniMicrosoft0.3282025-02-01evidence
Phi 4 ReasoningMicrosoft0.7332025-04-30evidence
Phi 4 Reasoning PlusMicrosoft0.792025-04-30evidence
Phi-3.5-mini-instructMicrosoft0.372024-08-23evidence
Phi-3.5-MoE-instructMicrosoft0.3792024-08-23evidence
Qwen2.5 72B InstructQwen0.8122024-09-19evidence
Qwen2.5 7B InstructQwen0.522024-09-19evidence
Qwen3 235B A22BQwen0.9562025-04-29evidence
Qwen3 30B A3BQwen0.912025-04-29evidence
Qwen3 32BQwen0.9382025-04-29evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.