Benchmark

NL2Repo

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-Flash-0731DeepSeek0.5422026-07-31evidence
DeepSeek-V4-Pro-0813DeepSeek0.6152026-08-13evidence
GLM-5.1Z.ai0.4272026-04-07evidence
GLM-5.2Z.ai0.4892026-06-16evidence
GLM-5.3Z.ai0.582026-08-14evidence
Hy3Tencent0.4562026-07-06evidence
MiniMax M2.7MiniMax0.3982026-03-18evidence
MiniMax M3MiniMax0.42132026-06-01evidence
Qwen3.6 PlusQwen0.3792026-04-02evidence
Qwen3.6-27BQwen0.3622026-04-21evidence
Qwen3.6-35B-A3BQwen0.2942026-04-16evidence
Qwen3.7 MaxQwen0.4722026-05-19evidence
Qwen3.7-PlusQwen0.4112026-05-31evidence
Qwen3.8 MaxQwen0.5592026-08-02evidence
Qwen3.8-27BQwen0.4232026-08-14evidence
Seed 2.1 ProByteDance0.472026-06-24evidence
Seed 2.1 TurboByteDance0.4372026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.