Benchmark

PostTrainBench

PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full…

Modality
text
Categories
reasoning, agents, code, systems
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5.2Z.ai0.3432026-06-16evidence
GLM-5.3Z.ai0.3982026-08-14evidence
Kimi K3Moonshot AI0.3662026-07-16evidence
MiniMax M3MiniMax0.3712026-06-01evidence
Seed 2.1 ProByteDance0.1652026-06-24evidence
Seed 2.1 TurboByteDance0.1832026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.