Benchmark
PinchBench
PinchBench evaluates coding agents on real-world agentic coding tasks, measuring both best-case and average performance across complex software…
- Modality
- text
- Categories
- agents, code
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 6
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| GLM-5V-Turbo | Z.ai | 0.807 | 2026-04-02 | evidence |
| LFM2.5-2.6B | Liquid AI | 0.6822 | 2026-08-04 | evidence |
| MiMo-V2-Omni | Xiaomi | 0.812 | 2026-03-18 | evidence |
| MiMo-V2-Pro | Xiaomi | 0.81 | 2026-03-18 | evidence |
| Nemotron 3 Ultra (550B A55B) | NVIDIA | 0.9 | 2026-06-04 | evidence |
| Nemotron 3.5 Lightning (30B A3B) | NVIDIA | 0.8537 | 2026-08-11 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.