Benchmark

FrontierCode

FrontierCode is Cognition's coding evaluation that tests whether models can pass difficult coding tasks while meeting the standards of high-quality…

Modality
text
Categories
reasoning, agents, code
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.4632026-06-09evidence
Claude Opus 5Anthropic0.5342026-07-24evidence
Claude Sonnet 5Anthropic0.3882026-06-30evidence
Grok 4.6xAI0.6132026-08-12evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.