Benchmark

SWE-Marathon

SWE-Marathon is an ultra-long-horizon software engineering benchmark covering tasks such as building compilers, optimizing kernels, and developing…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5.2Z.ai0.132026-06-16evidence
GLM-5.3Z.ai0.4252026-08-14evidence
Grok 4.5xAI0.292026-07-16evidence
Kimi K3Moonshot AI0.422026-07-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.