Benchmark

SWE-bench Pro

Harness and split version must be recorded with any score.

Released
2025-09-23
Categories
coding_agent
Openness
unknown
Source
model_reports
Reported scores
10

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude Opus 5Anthropic79.22026-07-24evidence
DeepSeek-V4-FlashDeepSeek52.62026-06-25evidence
DeepSeek-V4-ProDeepSeek52.12026-04-22evidence
DeepSeek-V4-ProDeepSeek54.42026-04-22evidence
DeepSeek-V4-ProDeepSeek55.42026-04-22evidence
Gemini 3.1 ProGoogle54.22026-02-19evidence
GLM-5.1Z.ai58.42026-04-03evidence
GLM-5.2Z.ai62.12026-06-16evidence
Grok 4.5xAI64.72026-07-08evidence
Hy4 previewTencent65.72026-08-28evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.