Benchmark

Wild Bench

WildBench is an automated evaluation framework that benchmarks large language models using 1,024 challenging, real-world tasks selected from over one…

Modality
text
Categories
reasoning, general, communication
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Jamba 1.5 LargeAI21 Labs0.4852024-08-22evidence
Jamba 1.5 MiniAI21 Labs0.4242024-08-22evidence
MiniStral 3 (14B Instruct 2512)Mistral0.6852025-12-04evidence
Ministral 3 (3B Instruct 2512)Mistral0.5682025-12-04evidence
Ministral 3 (8B Instruct 2512)Mistral0.6682025-12-04evidence
Mistral Large 3Mistral0.6852025-09-01evidence
Mistral Small 3 24B InstructMistral0.5222025-01-30evidence
Mistral Small 3.2 24B InstructMistral0.65332025-06-20evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.