Benchmark
MathVision
Prompt-format sensitive: vendors report boxed and unboxed variants and sometimes take the higher of the two.
- Released
- 2024-02-22
- Categories
- multimodal
- Openness
- unknown
- Source
- model_reports
- Reported scores
- 4
Reported scores
model_reports
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Gemma 4 (31B) | 85.6 | 2026-03-11 | evidence | |
| Kimi K3 | Moonshot AI | 97.8 | 2026-06-13 | evidence |
| Kimi K3 | Moonshot AI | 94.3 | 2026-06-13 | evidence |
| Qwen3.5-397B-A17B | Qwen | 88.6 | 2026-02-16 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.