Benchmark

CoWorkBench

CoWorkBench is Qwen's internal cowork benchmark for evaluating long-horizon office and productivity agent tasks across domains such as computer science…

Modality
text
Categories
productivity, reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3.7 MaxQwen0.6722026-05-19evidence
Qwen3.7-PlusQwen0.6512026-05-31evidence
Qwen3.8 MaxQwen0.7482026-08-02evidence
Qwen3.8-27BQwen0.7072026-08-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.