Benchmark

SkillsBench

SkillsBench evaluates coding agents on self-contained programming tasks, measuring practical engineering skills across diverse software development…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Hy3Tencent0.5532026-07-06evidence
Muse Glimmer-30BMeta0.4432026-08-10evidence
Qwen3.6 PlusQwen0.4572026-04-02evidence
Qwen3.6-27BQwen0.4822026-04-21evidence
Qwen3.6-35B-A3BQwen0.2872026-04-16evidence
Qwen3.7 MaxQwen0.5922026-05-19evidence
Qwen3.7-PlusQwen0.5492026-05-31evidence
Qwen3.8 MaxQwen0.7022026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.