Benchmark
Wild Bench
WildBench is an automated evaluation framework that benchmarks large language models using 1,024 challenging, real-world tasks selected from over one…
- Modality
- text
- Categories
- reasoning, general, communication
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 8
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Jamba 1.5 Large | AI21 Labs | 0.485 | 2024-08-22 | evidence |
| Jamba 1.5 Mini | AI21 Labs | 0.424 | 2024-08-22 | evidence |
| MiniStral 3 (14B Instruct 2512) | Mistral | 0.685 | 2025-12-04 | evidence |
| Ministral 3 (3B Instruct 2512) | Mistral | 0.568 | 2025-12-04 | evidence |
| Ministral 3 (8B Instruct 2512) | Mistral | 0.668 | 2025-12-04 | evidence |
| Mistral Large 3 | Mistral | 0.685 | 2025-09-01 | evidence |
| Mistral Small 3 24B Instruct | Mistral | 0.522 | 2025-01-30 | evidence |
| Mistral Small 3.2 24B Instruct | Mistral | 0.6533 | 2025-06-20 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.