Benchmark

TextVQA

TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Introduced to benchmark VQA models' ability to read and…

Modality
multimodal
Categories
multimodal, image_to_text, vision
Openness
unknown
Source
llm_stats
Reported scores
16

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek VL2DeepSeek0.8422024-12-13evidence
DeepSeek VL2 SmallDeepSeek0.8342024-12-13evidence
DeepSeek VL2 TinyDeepSeek0.8072024-12-13evidence
Gemma 3 12BGoogle0.6772025-03-12evidence
Gemma 3 27BGoogle0.6512025-03-12evidence
Gemma 3 4BGoogle0.5782025-03-12evidence
Grok-1.5VxAI0.7812024-04-12evidence
LFM2.5-VL-3BLiquid AI0.8432026-08-12evidence
Llama 3.2 90B InstructMeta0.7352024-09-25evidence
Nova LiteAmazon0.8022024-11-20evidence
Nova ProAmazon0.8152024-11-20evidence
Phi-3.5-vision-instructMicrosoft0.722024-08-23evidence
Phi-4-multimodal-instructMicrosoft0.7562025-02-01evidence
Qwen2-VL-72B-InstructQwen0.8552024-08-29evidence
Qwen2.5 VL 7B InstructQwen0.8492025-01-26evidence
Qwen2.5-Omni-7BQwen0.8442025-03-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.