Benchmark

Multi-SWE-Bench

What is Multi-SWE-Bench?

Multilingual SWE-bench variant spanning 7 programming languages; cross-repository context transfer is the measured capability.

Evaluates
coding agent
Openness
unknown
Catalog source
Model reports
Reported scores
2

Multi-SWE-Bench code

Multi-SWE-Bench results and reported scores

Model reports

ModelOrganizationReported valueReportedEvidence
Seed2.0 LiteByteDance41.12026-02-14evidence
Seed2.0 ProByteDance45.22026-02-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.

Which sources cite Multi-SWE-Bench?

Related coding agent benchmarks