Benchmark
ARC-AGI-3
Interactive multi-step format, so the agent harness is part of the measurement. Scores are low and spread wide, which makes small absolute gaps look…
- Released
- 2026-01-15
- Categories
- reasoning
- Openness
- unknown
- Source
- model_reports
- Reported scores
- 1
Reported scores
model_reports
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 30.16 | 2026-07-24 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.