Benchmark

EgoSchema

A diagnostic benchmark for very long-form video language understanding consisting of over 5000 human curated multiple choice questions based on 3-minute…

Modality
video
Categories
long_context, reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 1.0 ProGoogle0.5572024-02-15evidence
Gemini 2.0 FlashGoogle0.7152024-12-01evidence
Gemini 2.0 Flash-LiteGoogle0.6722025-02-05evidence
GPT-4oOpenAI0.7222024-08-06evidence
Nova LiteAmazon0.7142024-11-20evidence
Nova ProAmazon0.7212024-11-20evidence
Qwen2-VL-72B-InstructQwen0.7792024-08-29evidence
Qwen2.5 VL 72B InstructQwen0.7622025-01-26evidence
Qwen2.5-Omni-7BQwen0.6862025-03-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.