Benchmark
Program Bench
Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200…
- Modality
- text
- Categories
- agents, code
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 6
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| GLM-5.2 | Z.ai | 0.637 | 2026-06-16 | evidence |
| GLM-5.3 | Z.ai | 0.19 | 2026-08-14 | evidence |
| Kimi K2.7 Code | Moonshot AI | 0.536 | 2026-06-12 | evidence |
| Kimi K3 | Moonshot AI | 0.778 | 2026-07-16 | evidence |
| Seed 2.1 Pro | ByteDance | 0.503 | 2026-06-24 | evidence |
| Seed 2.1 Turbo | ByteDance | 0.494 | 2026-06-24 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.