Benchmark

Multi-SWE-Bench

A multilingual benchmark for issue resolving that evaluates Large Language Models' ability to resolve software issues across diverse programming…

Modality
text
Categories
reasoning, code
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K2-Thinking-0905Moonshot AI0.4192025-09-05evidence
MiniMax M2MiniMax0.3622025-10-27evidence
MiniMax M2.1MiniMax0.4942025-12-23evidence
MiniMax M2.5MiniMax0.5132026-02-12evidence
MiniMax M2.7MiniMax0.5272026-03-18evidence
Qwen3-Coder 480B A35B InstructQwen0.2582025-01-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.