Benchmark

MMMU (val)

Validation set of the Massive Multi-discipline Multimodal Understanding and Reasoning benchmark. Features college-level multimodal questions across 6 core…

Modality
multimodal
Categories
multimodal, reasoning, general, healthcare, vision
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemma 3 12BGoogle0.5962025-03-12evidence
Gemma 3 27BGoogle0.6492025-03-12evidence
Gemma 3 4BGoogle0.4882025-03-12evidence
LFM2.5-VL-3BLiquid AI0.4842026-08-12evidence
North Micro Vision InstructCohere0.3292026-08-12evidence
Qwen3 VL 30B A3B InstructQwen0.7422025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.762025-09-22evidence
Qwen3 VL 32B InstructQwen0.762025-09-22evidence
Qwen3 VL 32B ThinkingQwen0.7812025-09-22evidence
Qwen3 VL 4B InstructQwen0.6742025-09-22evidence
Qwen3 VL 4B ThinkingQwen0.7082025-09-22evidence
Qwen3 VL 8B InstructQwen0.6962025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.7412025-09-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.