Benchmark

We-Math

We-Math evaluates multimodal models on visual mathematical reasoning, requiring models to understand and solve math problems presented with visual…

Modality
multimodal
Categories
math, reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3.6 PlusQwen0.892026-04-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.