Benchmark
ChartQA
ChartQA is a large-scale benchmark comprising 9.6K human-written questions and 23.1K questions generated from human-written chart summaries, designed to…
- Modality
- multimodal
- Categories
- multimodal, reasoning, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 26
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Claude 3.5 Sonnet | Anthropic | 0.908 | 2024-10-22 | evidence |
| DeepSeek VL2 | DeepSeek | 0.86 | 2024-12-13 | evidence |
| DeepSeek VL2 Small | DeepSeek | 0.845 | 2024-12-13 | evidence |
| DeepSeek VL2 Tiny | DeepSeek | 0.81 | 2024-12-13 | evidence |
| Gemma 3 12B | 0.757 | 2025-03-12 | evidence | |
| Gemma 3 27B | 0.78 | 2025-03-12 | evidence | |
| Gemma 3 4B | 0.688 | 2025-03-12 | evidence | |
| GPT-4o | OpenAI | 0.857 | 2024-08-06 | evidence |
| Grok-1.5V | xAI | 0.761 | 2024-04-12 | evidence |
| LFM2.5-VL-3B | Liquid AI | 0.813 | 2026-08-12 | evidence |
| Llama 3.2 11B Instruct | Meta | 0.834 | 2024-09-25 | evidence |
| Llama 3.2 90B Instruct | Meta | 0.855 | 2024-09-25 | evidence |
| Llama 4 Maverick | Meta | 0.9 | 2025-04-05 | evidence |
| Llama 4 Scout | Meta | 0.888 | 2025-04-05 | evidence |
| Mistral Small 3.2 24B Instruct | Mistral | 0.874 | 2025-06-20 | evidence |
| North Micro Vision Instruct | Cohere | 0.808 | 2026-08-12 | evidence |
| Nova Lite | Amazon | 0.868 | 2024-11-20 | evidence |
| Nova Pro | Amazon | 0.892 | 2024-11-20 | evidence |
| Phi-3.5-vision-instruct | Microsoft | 0.818 | 2024-08-23 | evidence |
| Phi-4-multimodal-instruct | Microsoft | 0.814 | 2025-02-01 | evidence |
| Pixtral Large | Mistral | 0.881 | 2024-11-18 | evidence |
| Pixtral-12B | Mistral | 0.818 | 2024-09-17 | evidence |
| Qwen2-VL-72B-Instruct | Qwen | 0.883 | 2024-08-29 | evidence |
| Qwen2.5 VL 72B Instruct | Qwen | 0.895 | 2025-01-26 | evidence |
| Qwen2.5 VL 7B Instruct | Qwen | 0.873 | 2025-01-26 | evidence |
| Qwen2.5-Omni-7B | Qwen | 0.853 | 2025-03-27 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.