Benchmark

SWE-bench Verified (Agentic Coding)

SWE-bench Verified is a human-filtered subset of 500 software engineering problems drawn from real GitHub issues across 12 popular Python repositories…

Modality
text
Categories
reasoning, code
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Sonnet 4.5Anthropic0.7722025-09-29evidence
Kimi K2 InstructMoonshot AI0.6582025-07-11evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.