Benchmark

PinchBench

PinchBench evaluates coding agents on real-world agentic coding tasks, measuring both best-case and average performance across complex software…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5V-TurboZ.ai0.8072026-04-02evidence
LFM2.5-2.6BLiquid AI0.68222026-08-04evidence
MiMo-V2-OmniXiaomi0.8122026-03-18evidence
MiMo-V2-ProXiaomi0.812026-03-18evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.92026-06-04evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.85372026-08-11evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.