Benchmark

MLVU

A comprehensive benchmark for multi-task long video understanding that evaluates multimodal large language models on videos ranging from 3 minutes to 2…

Modality
multimodal
Categories
long_context, multimodal, video, vision
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen2.5 VL 7B InstructQwen0.7022025-01-26evidence
Qwen3 VL 235B A22B InstructQwen0.8432025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen0.8382025-09-22evidence
Qwen3.5-122B-A10BQwen0.8732026-02-24evidence
Qwen3.5-27BQwen0.8592026-02-24evidence
Qwen3.5-35B-A3BQwen0.8562026-02-24evidence
Qwen3.6 PlusQwen0.8672026-04-02evidence
Qwen3.6-27BQwen0.8662026-04-21evidence
Qwen3.6-35B-A3BQwen0.8622026-04-16evidence
Qwen3.7-PlusQwen0.8742026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.