Benchmark

MMMU

Image preprocessing and chain-of-thought settings affect results.

Released
2023-11-27
Categories
multimodal
Openness
unknown
Source
model_reports
Reported scores
4

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude Opus 4Anthropic76.52025-05-22evidence
Claude Sonnet 4Anthropic74.42025-05-22evidence
Gemini 2.5 ProGoogle82.02025-06-17evidence
Qwen3.5-397B-A17BQwen85.02026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.