Benchmark

MMAU

A massive multi-task audio understanding and reasoning benchmark comprising 10,000 carefully curated audio clips paired with human-annotated natural…

Modality
multimodal
Categories
multimodal, reasoning, audio
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Inkling-SmallThinking Machines Lab0.772026-07-30evidence
Nova 2 OmniAmazon0.7532025-12-02evidence
Qwen2.5-Omni-7BQwen0.6562025-03-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.