Benchmark

PostTrainBench

Measures a model's ability to run post-training itself. Hardware differs between reported runs (H100 vs H20), which moves the score.

Released
2026-01-20
Categories
ai_research
Openness
unknown
Source
model_reports
Reported scores
3

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
GLM-5.2Z.ai34.32026-06-16evidence
Hy4 previewTencent35.62026-08-28evidence
Kimi K3Moonshot AI36.62026-06-13evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.