Benchmark
EgoSchema
A diagnostic benchmark for very long-form video language understanding consisting of over 5000 human curated multiple choice questions based on 3-minute…
- Modality
- video
- Categories
- long_context, reasoning, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 9
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Gemini 1.0 Pro | 0.557 | 2024-02-15 | evidence | |
| Gemini 2.0 Flash | 0.715 | 2024-12-01 | evidence | |
| Gemini 2.0 Flash-Lite | 0.672 | 2025-02-05 | evidence | |
| GPT-4o | OpenAI | 0.722 | 2024-08-06 | evidence |
| Nova Lite | Amazon | 0.714 | 2024-11-20 | evidence |
| Nova Pro | Amazon | 0.721 | 2024-11-20 | evidence |
| Qwen2-VL-72B-Instruct | Qwen | 0.779 | 2024-08-29 | evidence |
| Qwen2.5 VL 72B Instruct | Qwen | 0.762 | 2025-01-26 | evidence |
| Qwen2.5-Omni-7B | Qwen | 0.686 | 2025-03-27 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.