Benchmark

MMLongBench-Doc

MMLongBench-Doc evaluates long document understanding capabilities in vision-language models.

Modality
image
Categories
long_context, multimodal, vision
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3 VL 235B A22B ThinkingQwen0.5622025-09-22evidence
Qwen3.5-122B-A10BQwen0.592026-02-24evidence
Qwen3.5-27BQwen0.6022026-02-24evidence
Qwen3.5-35B-A3BQwen0.5952026-02-24evidence
Qwen3.6 PlusQwen0.622026-04-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.