Benchmark

Seal-0

Seal-0 is a benchmark for evaluating agentic search capabilities, testing models' ability to navigate and retrieve information using tools.

Modality
text
Categories
reasoning, search
Openness
unknown
Source
llm_stats
Reported scores
6

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K2-Thinking-0905Moonshot AI0.5632025-09-05evidence
Kimi K2.5Moonshot AI0.5742026-01-27evidence
Qwen3.5-122B-A10BQwen0.4412026-02-24evidence
Qwen3.5-27BQwen0.4722026-02-24evidence
Qwen3.5-35B-A3BQwen0.4142026-02-24evidence
Qwen3.5-397B-A17BQwen0.4692026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.