Benchmark

CL-bench

CL-bench is an open-source benchmark with its own data and rubrics for evaluating models on coding and agentic tasks, scored using a setup fully aligned…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Hy3Tencent0.2382026-07-06evidence
MiniMax M3MiniMax0.20482026-06-01evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.