Benchmark

DeepSearchQA

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and…

Modality
text
Categories
reasoning, search, agents
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.6Anthropic0.9132026-02-05evidence
Claude Opus 4.8Anthropic0.9312026-05-28evidence
Hy3Tencent0.912026-07-06evidence
Kimi K2.5Moonshot AI0.7712026-01-27evidence
Kimi K2.6Moonshot AI0.832026-04-20evidence
Kimi K3Moonshot AI0.952026-07-16evidence
MiMo-V2-ProXiaomi0.8672026-03-18evidence
Muse Glimmer-30BMeta0.7462026-08-10evidence
Muse SparkMeta0.7482026-04-08evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.