Benchmark

SWE-bench Multilingual

Per-language resolution rates differ widely, so an average hides which languages the model actually handles.

Released
2025-04-22
Categories
coding_agent
Openness
unknown
Source
model_reports
Reported scores
7

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-FlashDeepSeek73.32026-06-25evidence
DeepSeek-V4-ProDeepSeek69.82026-04-22evidence
DeepSeek-V4-ProDeepSeek74.12026-04-22evidence
DeepSeek-V4-ProDeepSeek76.22026-04-22evidence
GLM-5Z.ai73.32026-02-11evidence
Hy4 previewTencent82.92026-08-28evidence
Qwen3.5-397B-A17BQwen69.32026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.