Benchmark

MT-Bench

MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging…

Modality
text
Categories
reasoning, roleplay, general, communication, creativity
Openness
unknown
Source
llm_stats
Reported scores
12

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V2.5DeepSeek0.9022024-05-08evidence
Hermes 3 70BNous Research8.992024-08-15evidence
Llama 3.1 Nemotron 70B InstructNVIDIA0.08992024-10-01evidence
Llama 3.1 Nemotron Nano 8B V1NVIDIA0.812025-03-18evidence
Llama-3.3 Nemotron Super 49B v1NVIDIA0.9172025-03-18evidence
Ministral 8B InstructMistral0.832024-10-16evidence
Mistral Large 2Mistral0.8632024-07-24evidence
Mistral Small 3 24B InstructMistral0.8352025-01-30evidence
Pixtral-12BMistral0.7682024-09-17evidence
Qwen2 7B InstructQwen0.8412024-07-23evidence
Qwen2.5 72B InstructQwen0.9352024-09-19evidence
Qwen2.5 7B InstructQwen0.8752024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.