Benchmark

MMMUval

Validation set for MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning) benchmark, designed to evaluate multimodal models on massive…

Modality
multimodal
Categories
multimodal, reasoning, general, healthcare, vision
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Sonnet 4.5Anthropic0.7782025-09-29evidence
Qwen2-VL-72B-InstructQwen0.6452024-08-29evidence
Qwen3 VL 235B A22B InstructQwen0.7872025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen0.8062025-09-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.