Benchmark

ARC-AGI

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract…

Modality
image
Categories
reasoning, spatial_reasoning, vision
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5.2OpenAI0.8622025-12-11evidence
GPT-5.2 ProOpenAI0.9052025-12-11evidence
GPT-5.4OpenAI0.9372026-03-05evidence
GPT-5.5OpenAI0.952026-04-23evidence
Inkling-SmallThinking Machines Lab0.842026-07-30evidence
LongCat-Flash-ThinkingMeituan0.5032025-09-22evidence
o3OpenAI0.882025-04-16evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.4182025-07-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.