Benchmark

SWE-Marathon Benchmark

What is SWE-Marathon?

SWE-Marathon evaluates AI agents on 20 ultra-long-horizon software engineering tasks, each with a unique executable environment, a human-written reference…

Released
2026-06-05
Evaluates
Code, Code & Software, Software & AI Compute
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

SWE-Marathon paper

No reported scores are on record for this benchmark yet.

Which sources cite SWE-Marathon?

Related Code, Code & Software, Software & AI Compute benchmarks