Benchmark

Program Bench

Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5.2Z.ai0.6372026-06-16evidence
GLM-5.3Z.ai0.192026-08-14evidence
Kimi K2.7 CodeMoonshot AI0.5362026-06-12evidence
Kimi K3Moonshot AI0.7782026-07-16evidence
Seed 2.1 ProByteDance0.5032026-06-24evidence
Seed 2.1 TurboByteDance0.4942026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.