Benchmark

VideoMME w sub.

The first-ever comprehensive evaluation benchmark of Multi-modal LLMs in Video analysis. Features 900 videos (254 hours) with 2,700 question-answer pairs…

Modality
multimodal
Categories
multimodal, video, vision
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5OpenAI0.8672025-08-07evidence
Qwen2.5 VL 32B InstructQwen0.7792025-02-28evidence
Qwen2.5 VL 7B InstructQwen0.7162025-01-26evidence
Qwen2.5-Omni-7BQwen0.7242025-03-27evidence
Qwen3.5-122B-A10BQwen0.8732026-02-24evidence
Qwen3.5-27BQwen0.872026-02-24evidence
Qwen3.5-35B-A3BQwen0.8662026-02-24evidence
Qwen3.6-27BQwen0.8772026-04-21evidence
Qwen3.6-35B-A3BQwen0.8662026-04-16evidence
Qwen3.8 MaxQwen0.9042026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.