Benchmark

DeepPlanning

DeepPlanning evaluates LLMs on complex multi-step planning tasks requiring long-horizon reasoning, goal decomposition, and strategic decision-making.

Modality
text
Categories
reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen3.5-122B-A10BQwen0.2412026-02-24evidence
Qwen3.5-27BQwen0.2262026-02-24evidence
Qwen3.5-35B-A3BQwen0.2282026-02-24evidence
Qwen3.5-397B-A17BQwen0.3432026-02-16evidence
Qwen3.5-4BQwen0.1762026-03-02evidence
Qwen3.5-9BQwen0.182026-03-02evidence
Qwen3.6 PlusQwen0.4152026-04-02evidence
Qwen3.6-35B-A3BQwen0.2592026-04-16evidence
Qwen3.7-PlusQwen0.6232026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.