Benchmark
TextVQA
TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Introduced to benchmark VQA models' ability to read and…
- Modality
- multimodal
- Categories
- multimodal, image_to_text, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 16
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| DeepSeek VL2 | DeepSeek | 0.842 | 2024-12-13 | evidence |
| DeepSeek VL2 Small | DeepSeek | 0.834 | 2024-12-13 | evidence |
| DeepSeek VL2 Tiny | DeepSeek | 0.807 | 2024-12-13 | evidence |
| Gemma 3 12B | 0.677 | 2025-03-12 | evidence | |
| Gemma 3 27B | 0.651 | 2025-03-12 | evidence | |
| Gemma 3 4B | 0.578 | 2025-03-12 | evidence | |
| Grok-1.5V | xAI | 0.781 | 2024-04-12 | evidence |
| LFM2.5-VL-3B | Liquid AI | 0.843 | 2026-08-12 | evidence |
| Llama 3.2 90B Instruct | Meta | 0.735 | 2024-09-25 | evidence |
| Nova Lite | Amazon | 0.802 | 2024-11-20 | evidence |
| Nova Pro | Amazon | 0.815 | 2024-11-20 | evidence |
| Phi-3.5-vision-instruct | Microsoft | 0.72 | 2024-08-23 | evidence |
| Phi-4-multimodal-instruct | Microsoft | 0.756 | 2025-02-01 | evidence |
| Qwen2-VL-72B-Instruct | Qwen | 0.855 | 2024-08-29 | evidence |
| Qwen2.5 VL 7B Instruct | Qwen | 0.849 | 2025-01-26 | evidence |
| Qwen2.5-Omni-7B | Qwen | 0.844 | 2025-03-27 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.