Benchmark

WildClawBench

WildClawBench is an agentic coding benchmark from InternLM/Claw-Eval that reports overall model performance on real-world tool-using development tasks.

Modality
text
Categories
agents, coding
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Hy3Tencent0.5362026-07-06evidence
MiMo-V2.5-ProXiaomi0.432026-04-27evidence
Muse Glimmer-30BMeta0.4762026-08-10evidence
Seed 2.1 ProByteDance0.6172026-06-24evidence
Seed 2.1 TurboByteDance0.6282026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.