Benchmark

SWE-bench Verified (Agentless)

A human-validated subset of SWE-bench that evaluates language models' ability to resolve real-world GitHub issues using an agentless approach. The…

Modality
text
Categories
reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K2 InstructMoonshot AI0.5182025-07-11evidence
MiMo-V2.5-ProXiaomi0.3572026-04-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.