Benchmark

TOMATO

TOMATO (Temporal Reasoning Multimodal Evaluation) assesses multimodal models on motion and temporal perception in video, testing understanding of actions…

Modality
multimodal
Categories
multimodal, reasoning, video, vision
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Seed 2.1 ProByteDance0.7952026-06-24evidence
Seed 2.1 TurboByteDance0.5682026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.