Benchmark

ARC-AGI v2

ARC-AGI-2 is an upgraded benchmark for measuring abstract reasoning and problem-solving abilities in AI systems through visual grid transformation tasks…

Modality
multimodal
Categories
reasoning, spatial_reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4Anthropic0.0862025-05-22evidence
Claude Opus 4.5Anthropic0.3762025-11-24evidence
Claude Opus 4.6Anthropic0.6882026-02-05evidence
Claude Sonnet 4.6Anthropic0.5832026-02-17evidence
Gemini 2.5 ProGoogle0.0492025-05-20evidence
Gemini 3 FlashGoogle0.3362025-12-17evidence
Gemini 3 ProGoogle0.3112025-11-18evidence
Gemini 3.1 ProGoogle0.7712026-02-19evidence
Gemini 3.5 FlashGoogle0.7212026-05-19evidence
GPT-5.2OpenAI0.5292025-12-11evidence
GPT-5.2 ProOpenAI0.5422025-12-11evidence
GPT-5.4OpenAI0.7332026-03-05evidence
GPT-5.5OpenAI0.852026-04-23evidence
Grok-4xAI0.1592025-07-09evidence
Inkling-SmallThinking Machines Lab0.4012026-07-30evidence
Muse SparkMeta0.4252026-04-08evidence
o3OpenAI0.0652025-04-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.