Benchmark

MM-MT-Bench

A multi-turn LLM-as-a-judge evaluation benchmark for testing multimodal instruction-tuned models' ability to follow user instructions in multi-turn…

Modality
multimodal
Categories
multimodal, communication
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
MiniStral 3 (14B Instruct 2512)Mistral0.08492025-12-04evidence
Ministral 3 (3B Instruct 2512)Mistral0.07832025-12-04evidence
Ministral 3 (8B Instruct 2512)Mistral0.08082025-12-04evidence
Mistral Large 3Mistral84.92025-09-01evidence
Pixtral LargeMistral0.742024-11-18evidence
Pixtral-12BMistral0.6052024-09-17evidence
Qwen2.5-Omni-7BQwen0.062025-03-27evidence
Qwen3 VL 235B A22B InstructQwen8.52025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen8.52025-09-22evidence
Qwen3 VL 30B A3B InstructQwen8.12025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen7.92025-09-22evidence
Qwen3 VL 32B InstructQwen8.42025-09-22evidence
Qwen3 VL 32B ThinkingQwen8.32025-09-22evidence
Qwen3 VL 4B InstructQwen7.52025-09-22evidence
Qwen3 VL 4B ThinkingQwen7.72025-09-22evidence
Qwen3 VL 8B InstructQwen7.72025-09-22evidence
Qwen3 VL 8B ThinkingQwen8.02025-09-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.