Benchmark

Video-MME

Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos…

Modality
multimodal
Categories
multimodal, reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 1.5 FlashGoogle0.7612024-05-01evidence
Gemini 1.5 Flash 8BGoogle0.6622024-03-15evidence
Gemini 1.5 ProGoogle0.7862024-05-01evidence
Gemini 2.5 ProGoogle0.8482025-05-20evidence
Kimi K2.5Moonshot AI0.8742026-01-27evidence
MiMo-V2.5Xiaomi0.8772026-04-22evidence
MiniMax M3MiniMax0.8542026-06-01evidence
Nova 2 OmniAmazon0.7792025-12-02evidence
Phi-4-multimodal-instructMicrosoft0.552025-02-01evidence
Qwen3 VL 30B A3B InstructQwen0.7452025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.7332025-09-22evidence
Qwen3 VL 8B InstructQwen0.7142025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.7182025-09-22evidence
Qwen3.6 PlusQwen0.8422026-04-02evidence
Qwen3.7-PlusQwen0.882026-05-31evidence
Seed 2.1 ProByteDance0.8922026-06-24evidence
Seed 2.1 TurboByteDance0.892026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.