Benchmark

ProgramBench

Cleanroom program-rebuild tasks; most models score low, so small differences sit within run-to-run noise.

Released
2026-05-05
Categories
coding
Openness
unknown
Source
model_reports
Reported scores
1

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Hy4 previewTencent17.52026-08-28evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.