Benchmark

SWE-bench Multilingual

A multilingual benchmark for issue resolving in software engineering that covers Java, TypeScript, JavaScript, Go, Rust, C, and C++. Contains 1,632…

Modality
text
Categories
reasoning, code
Openness
unknown
Source
llm_stats
Reported scores
38

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Mythos PreviewAnthropic0.8732026-04-07evidence
Claude Opus 4.6Anthropic0.77832026-02-05evidence
Claude Opus 4.8Anthropic0.8442026-05-28evidence
Claude Sonnet 5Anthropic0.7832026-06-30evidence
DeepSeek-R1-0528DeepSeek0.3052025-05-28evidence
DeepSeek-V3.1DeepSeek0.5452025-01-10evidence
DeepSeek-V3.2DeepSeek0.7022025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.7022025-12-01evidence
DeepSeek-V3.2-ExpDeepSeek0.5792025-09-29evidence
DeepSeek-V4-Flash-0423DeepSeek0.7022026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.7332026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.7622026-04-23evidence
GLM-4.7Z.ai0.6672025-12-22evidence
Hy3Tencent0.7582026-07-06evidence
Kimi K2 InstructMoonshot AI0.4732025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.4732025-09-05evidence
Kimi K2-Thinking-0905Moonshot AI0.6112025-09-05evidence
Kimi K2.5Moonshot AI0.732026-01-27evidence
Kimi K2.6Moonshot AI0.7672026-04-20evidence
Laguna S 2.1Poolside0.7852026-07-21evidence
Laguna XS 2.1Poolside0.6312026-07-02evidence
LongCat-Flash-LiteMeituan0.3812026-02-05evidence
MAI-Code-1-FlashMicrosoft0.6552026-06-02evidence
MiMo-V2-FlashXiaomi0.7172025-12-16evidence
MiMo-V2-ProXiaomi0.7172026-03-18evidence
MiniMax M2MiniMax0.5652025-10-27evidence
MiniMax M2.1MiniMax0.7252025-12-23evidence
MiniMax M2.7MiniMax0.7652026-03-18evidence
Nemotron 3 Super (120B A12B)NVIDIA0.45782026-03-11evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.6772026-06-04evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.39332026-08-11evidence
Qwen3-Coder 480B A35B InstructQwen0.5472025-01-31evidence
Qwen3.5-397B-A17BQwen0.6932026-02-16evidence
Qwen3.6 PlusQwen0.7382026-04-02evidence
Qwen3.6-27BQwen0.7132026-04-21evidence
Qwen3.6-35B-A3BQwen0.6722026-04-16evidence
Qwen3.7 MaxQwen0.7832026-05-19evidence
Qwen3.7-PlusQwen0.7582026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.