Benchmark

DocVQA

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's…

Modality
multimodal
Categories
multimodal, image_to_text, vision
Openness
unknown
Source
llm_stats
Reported scores
28

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3.5 SonnetAnthropic0.9522024-10-22evidence
DeepSeek VL2DeepSeek0.9332024-12-13evidence
DeepSeek VL2 SmallDeepSeek0.9232024-12-13evidence
DeepSeek VL2 TinyDeepSeek0.8892024-12-13evidence
Gemma 3 12BGoogle0.8712025-03-12evidence
Gemma 3 27BGoogle0.8662025-03-12evidence
Gemma 3 4BGoogle0.7582025-03-12evidence
GPT-4oOpenAI0.9282024-08-06evidence
Grok-1.5xAI0.8562024-03-28evidence
Grok-1.5VxAI0.8562024-04-12evidence
Grok-2xAI0.9362024-08-13evidence
Grok-2 minixAI0.9322024-08-13evidence
LFM2.5-VL-3BLiquid AI0.9112026-08-12evidence
Llama 3.2 11B InstructMeta0.8842024-09-25evidence
Llama 3.2 90B InstructMeta0.9012024-09-25evidence
Llama 4 MaverickMeta0.9442025-04-05evidence
Llama 4 ScoutMeta0.9442025-04-05evidence
Mistral Small 3.2 24B InstructMistral0.94862025-06-20evidence
North Micro Vision InstructCohere0.9212026-08-12evidence
Nova LiteAmazon0.9242024-11-20evidence
Nova ProAmazon0.9352024-11-20evidence
Phi-4-multimodal-instructMicrosoft0.9322025-02-01evidence
Pixtral LargeMistral0.9332024-11-18evidence
Pixtral-12BMistral0.9072024-09-17evidence
Qwen2.5 VL 32B InstructQwen0.9482025-02-28evidence
Qwen2.5 VL 72B InstructQwen0.9642025-01-26evidence
Qwen2.5 VL 7B InstructQwen0.9572025-01-26evidence
Qwen2.5-Omni-7BQwen0.9522025-03-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.