Benchmark

ARC-AGI-3

ARC-AGI-3 is the third-generation Abstraction and Reasoning Corpus benchmark, an interactive-reasoning evaluation designed to measure fluid, novel…

Modality
multimodal
Categories
reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 5Anthropic0.3022026-07-24evidence
GPT-5.6 LunaOpenAI0.00182026-07-09evidence
GPT-5.6 SolOpenAI0.07782026-07-09evidence
GPT-5.6 TerraOpenAI0.0082026-07-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.