Benchmark

Arena-Hard v2

Arena-Hard-Auto v2 is a challenging benchmark consisting of 500 carefully curated prompts sourced from Chatbot Arena and WildChat-1M, designed to evaluate…

Modality
text
Categories
reasoning, general, creativity, writing
Openness
unknown
Source
llm_stats
Reported scores
16

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
MiMo-V2-FlashXiaomi0.8622025-12-16evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.6772025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.73882026-03-11evidence
Qwen3 VL 235B A22B InstructQwen0.7742025-09-22evidence
Qwen3 VL 30B A3B InstructQwen0.5852025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.5672025-09-22evidence
Qwen3 VL 32B InstructQwen0.6472025-09-22evidence
Qwen3 VL 32B ThinkingQwen0.6052025-09-22evidence
Qwen3 VL 4B ThinkingQwen0.3682025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.5112025-09-22evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.7922025-07-22evidence
Qwen3-235B-A22B-Thinking-2507Qwen0.7972025-07-25evidence
Qwen3-Next-80B-A3B-InstructQwen0.8272025-09-10evidence
Qwen3-Next-80B-A3B-ThinkingQwen0.6232025-09-10evidence
Sarvam-105BSarvam AI0.712026-03-06evidence
Sarvam-30BSarvam AI0.492026-03-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.