Benchmark

GraphWalks

GraphWalks is a synthetic multi-hop long-context reasoning benchmark in which a model is given an edge-list representation of a graph and must traverse it…

Modality
text
Categories
long_context, reasoning
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
MAI-Thinking-1Microsoft0.92026-06-02evidence
MiMo-V2.5Xiaomi0.872026-04-22evidence
MiMo-V2.5-ProXiaomi0.622026-04-27evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.