Benchmark

FrontierSWE

FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
16

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.92026-06-09evidence
Claude Opus 4.6Anthropic0.562026-02-05evidence
Claude Opus 4.7Anthropic0.632026-04-16evidence
Claude Opus 4.8Anthropic0.752026-05-28evidence
DeepSeek-V4-Pro-MaxDeepSeek0.292026-04-23evidence
Gemini 3.1 ProGoogle0.42026-02-19evidence
GLM-5.1Z.ai0.312026-04-07evidence
GLM-5.2Z.ai0.742026-06-16evidence
GLM-5.3Z.ai0.7812026-08-14evidence
GPT-5.4OpenAI0.542026-03-05evidence
GPT-5.5OpenAI0.732026-04-23evidence
Kimi K2.5Moonshot AI0.262026-01-27evidence
Kimi K2.6Moonshot AI0.272026-04-20evidence
Kimi K3Moonshot AI0.8122026-07-16evidence
Qwen3.6 PlusQwen0.222026-04-02evidence
Qwen3.8 MaxQwen0.7352026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.