Benchmark

WideSearch

WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests…

Modality
text
Categories
reasoning, search, agents
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Hy3Tencent0.7642026-07-06evidence
Kimi K2.5Moonshot AI0.792026-01-27evidence
Kimi K2.6Moonshot AI0.8082026-04-20evidence
Qwen3.5-122B-A10BQwen0.6052026-02-24evidence
Qwen3.5-27BQwen0.6112026-02-24evidence
Qwen3.5-35B-A3BQwen0.5712026-02-24evidence
Qwen3.5-397B-A17BQwen0.742026-02-16evidence
Qwen3.6 PlusQwen0.7432026-04-02evidence
Qwen3.6-35B-A3BQwen0.6012026-04-16evidence
Qwen3.8 MaxQwen0.8192026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.