Benchmark

RealWorldQA

RealWorldQA is a benchmark designed to evaluate basic real-world spatial understanding capabilities of multimodal models. The initial release consists of…

Modality
multimodal
Categories
spatial_reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
29

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek VL2DeepSeek0.6842024-12-13evidence
DeepSeek VL2 SmallDeepSeek0.6542024-12-13evidence
DeepSeek VL2 TinyDeepSeek0.6422024-12-13evidence
Grok-1.5VxAI0.6872024-04-12evidence
LFM2.5-VL-3BLiquid AI0.7312026-08-12evidence
North Micro Vision InstructCohere0.6222026-08-12evidence
Qwen2-VL-72B-InstructQwen0.7782024-08-29evidence
Qwen2.5-Omni-7BQwen0.7032025-03-27evidence
Qwen3 VL 235B A22B InstructQwen0.7932025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen0.8132025-09-22evidence
Qwen3 VL 30B A3B InstructQwen0.7372025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.7742025-09-22evidence
Qwen3 VL 32B InstructQwen0.792025-09-22evidence
Qwen3 VL 32B ThinkingQwen0.7842025-09-22evidence
Qwen3 VL 4B InstructQwen0.7092025-09-22evidence
Qwen3 VL 4B ThinkingQwen0.7322025-09-22evidence
Qwen3 VL 8B InstructQwen0.7152025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.7352025-09-22evidence
Qwen3.5-122B-A10BQwen0.8512026-02-24evidence
Qwen3.5-27BQwen0.8372026-02-24evidence
Qwen3.5-35B-A3BQwen0.8412026-02-24evidence
Qwen3.6 PlusQwen0.8542026-04-02evidence
Qwen3.6-27BQwen0.8412026-04-21evidence
Qwen3.6-35B-A3BQwen0.8532026-04-16evidence
Qwen3.7-PlusQwen0.8692026-05-31evidence
Qwen3.8 MaxQwen0.882026-08-02evidence
Qwen3.8-27BQwen0.8592026-08-14evidence
Seed 2.1 ProByteDance0.8672026-06-24evidence
Seed 2.1 TurboByteDance0.8632026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.