Benchmark

MathVision

Prompt-format sensitive: vendors report boxed and unboxed variants and sometimes take the higher of the two.

Released
2024-02-22
Categories
multimodal
Openness
unknown
Source
model_reports
Reported scores
4

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Gemma 4 (31B)Google85.62026-03-11evidence
Kimi K3Moonshot AI97.82026-06-13evidence
Kimi K3Moonshot AI94.32026-06-13evidence
Qwen3.5-397B-A17BQwen88.62026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.