Benchmark

QwenWorldBench

QwenWorldBench is Qwen's internal benchmark for evaluating LLMs as world models that simulate agentic environments across Terminal, SWE, MCP, Search, OS…

Modality
text
Categories
reasoning, simulation, agents
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3.7 MaxQwen0.5732026-05-19evidence
Qwen3.7-PlusQwen0.6212026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.