Benchmark

ChartQA

ChartQA is a large-scale benchmark comprising 9.6K human-written questions and 23.1K questions generated from human-written chart summaries, designed to…

Modality
multimodal
Categories
multimodal, reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
26

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3.5 SonnetAnthropic0.9082024-10-22evidence
DeepSeek VL2DeepSeek0.862024-12-13evidence
DeepSeek VL2 SmallDeepSeek0.8452024-12-13evidence
DeepSeek VL2 TinyDeepSeek0.812024-12-13evidence
Gemma 3 12BGoogle0.7572025-03-12evidence
Gemma 3 27BGoogle0.782025-03-12evidence
Gemma 3 4BGoogle0.6882025-03-12evidence
GPT-4oOpenAI0.8572024-08-06evidence
Grok-1.5VxAI0.7612024-04-12evidence
LFM2.5-VL-3BLiquid AI0.8132026-08-12evidence
Llama 3.2 11B InstructMeta0.8342024-09-25evidence
Llama 3.2 90B InstructMeta0.8552024-09-25evidence
Llama 4 MaverickMeta0.92025-04-05evidence
Llama 4 ScoutMeta0.8882025-04-05evidence
Mistral Small 3.2 24B InstructMistral0.8742025-06-20evidence
North Micro Vision InstructCohere0.8082026-08-12evidence
Nova LiteAmazon0.8682024-11-20evidence
Nova ProAmazon0.8922024-11-20evidence
Phi-3.5-vision-instructMicrosoft0.8182024-08-23evidence
Phi-4-multimodal-instructMicrosoft0.8142025-02-01evidence
Pixtral LargeMistral0.8812024-11-18evidence
Pixtral-12BMistral0.8182024-09-17evidence
Qwen2-VL-72B-InstructQwen0.8832024-08-29evidence
Qwen2.5 VL 72B InstructQwen0.8952025-01-26evidence
Qwen2.5 VL 7B InstructQwen0.8732025-01-26evidence
Qwen2.5-Omni-7BQwen0.8532025-03-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.