Benchmark

InfoVQA

InfoVQA dataset with 30,000 questions and 5,000 infographic images requiring joint reasoning over document layout, textual content, graphical elements…

Modality
multimodal
Categories
multimodal, vision
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek VL2DeepSeek0.7812024-12-13evidence
DeepSeek VL2 SmallDeepSeek0.7582024-12-13evidence
DeepSeek VL2 TinyDeepSeek0.6612024-12-13evidence
Gemma 3 12BGoogle0.6492025-03-12evidence
Gemma 3 27BGoogle0.7062025-03-12evidence
Gemma 3 4BGoogle0.52025-03-12evidence
North Micro Vision InstructCohere0.6522026-08-12evidence
Phi-4-multimodal-instructMicrosoft0.7272025-02-01evidence
Qwen2.5 VL 32B InstructQwen0.8342025-02-28evidence
Qwen2.5 VL 7B InstructQwen0.8262025-01-26evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.