Benchmark

FrontierCode 1.1

FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for…

Modality
text
Categories
reasoning, agents, code
Openness
unknown
Source
llm_stats
Reported scores
16

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.5352026-06-09evidence
Claude Opus 4.7Anthropic0.3852026-04-16evidence
Claude Opus 4.8Anthropic0.4652026-05-28evidence
Claude Opus 5Anthropic0.5342026-07-24evidence
Claude Sonnet 5Anthropic0.4272026-06-30evidence
DeepSeek-V4-Pro-MaxDeepSeek0.1762026-04-23evidence
Gemini 3.7 FlashGoogle0.4362026-08-13evidence
GLM-5.2Z.ai0.2452026-06-16evidence
GPT-5.5OpenAI0.432026-04-23evidence
GPT-5.6 LunaOpenAI0.3982026-07-09evidence
GPT-5.6 SolOpenAI0.4752026-07-09evidence
GPT-5.6 TerraOpenAI0.4132026-07-09evidence
Grok 4.5xAI0.4242026-07-16evidence
Kimi K2.7 CodeMoonshot AI0.3012026-06-12evidence
MiniMax M3MiniMax0.1472026-06-01evidence
Qwen3.7-PlusQwen0.1022026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.