Benchmark

SWE-bench Verified

Scaffold, tool permissions and time limit are part of the result; vendors run their own harnesses.

Released
2024-08-13
Categories
coding_agent
Openness
unknown
Source
model_reports
Reported scores
16

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude 3.7 SonnetAnthropic70.32025-02-24evidence
Claude 3.7 SonnetAnthropic63.72025-02-24evidence
Claude Opus 4Anthropic72.52025-05-22evidence
Claude Opus 5Anthropic96.02026-07-24evidence
Claude Sonnet 4Anthropic72.72025-05-22evidence
DeepSeek-R1DeepSeek49.22025-01-22evidence
DeepSeek-V3DeepSeek42.02024-12-27evidence
DeepSeek-V4-FlashDeepSeek79.02026-06-25evidence
DeepSeek-V4-ProDeepSeek73.62026-04-22evidence
DeepSeek-V4-ProDeepSeek79.42026-04-22evidence
DeepSeek-V4-ProDeepSeek80.62026-04-22evidence
Gemini 2.5 ProGoogle67.22025-06-17evidence
Gemini 2.5 ProGoogle59.62025-06-17evidence
Gemini 3.1 ProGoogle80.62026-02-19evidence
GLM-5Z.ai77.82026-02-11evidence
Qwen3.5-397B-A17BQwen76.42026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.