Benchmark

BrowseComp-zh

A high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web, consisting of 289 multi-hop questions spanning 11…

Modality
text
Categories
reasoning, search
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1-0528DeepSeek0.3572025-05-28evidence
DeepSeek-V3.1DeepSeek0.4922025-01-10evidence
DeepSeek-V3.2DeepSeek0.652025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.652025-12-01evidence
DeepSeek-V3.2-ExpDeepSeek0.4792025-09-29evidence
GLM-4.7Z.ai0.6662025-12-22evidence
Kimi K2-Thinking-0905Moonshot AI0.6232025-09-05evidence
LongCat-Flash-Thinking-2601Meituan0.692026-01-14evidence
MiniMax M2MiniMax0.4852025-10-27evidence
Qwen3.5-122B-A10BQwen0.6992026-02-24evidence
Qwen3.5-27BQwen0.6212026-02-24evidence
Qwen3.5-35B-A3BQwen0.6952026-02-24evidence
Qwen3.5-397B-A17BQwen0.7032026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.