Benchmark

Claw-Eval

Claw-Eval tests real-world agentic task completion across complex multi-step scenarios, evaluating a model's ability to use tools, navigate environments…

Modality
text
Categories
agents, code
Openness
unknown
Source
llm_stats
Reported scores
14

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5V-TurboZ.ai0.752026-04-02evidence
Hy3Tencent0.6852026-07-06evidence
Kimi K2.6Moonshot AI0.8092026-04-20evidence
LFM2.5-2.6BLiquid AI0.62852026-08-04evidence
MiMo-V2-OmniXiaomi0.5482026-03-18evidence
MiMo-V2-ProXiaomi0.6152026-03-18evidence
MiMo-V2.5Xiaomi0.6322026-04-22evidence
MiMo-V2.5-ProXiaomi0.642026-04-27evidence
MiniMax M3MiniMax0.7452026-06-01evidence
Qwen3.6 PlusQwen0.5872026-04-02evidence
Qwen3.6-27BQwen0.6062026-04-21evidence
Qwen3.6-35B-A3BQwen0.52026-04-16evidence
Qwen3.7 MaxQwen0.6522026-05-19evidence
Qwen3.7-PlusQwen0.6272026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.