Benchmark

BLINK

BLINK: Multimodal Large Language Models Can See but Not Perceive. A benchmark for multimodal language models focusing on core visual perception abilities…

Modality
multimodal
Categories
multimodal, reasoning, spatial_reasoning, 3d, vision
Openness
unknown
Source
llm_stats
Reported scores
15

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
LFM2.5-VL-3BLiquid AI0.6152026-08-12evidence
North Micro Vision InstructCohere0.5272026-08-12evidence
Phi-4-multimodal-instructMicrosoft0.6132025-02-01evidence
Qwen3 VL 235B A22B InstructQwen0.7072025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen0.6712025-09-22evidence
Qwen3 VL 30B A3B InstructQwen0.6772025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.6542025-09-22evidence
Qwen3 VL 32B InstructQwen0.6732025-09-22evidence
Qwen3 VL 32B ThinkingQwen0.6852025-09-22evidence
Qwen3 VL 4B InstructQwen0.6582025-09-22evidence
Qwen3 VL 4B ThinkingQwen0.6342025-09-22evidence
Qwen3 VL 8B InstructQwen0.6912025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.6872025-09-22evidence
Seed 2.1 ProByteDance0.8142026-06-24evidence
Seed 2.1 TurboByteDance0.7942026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.