Benchmark

Graphwalks parents >128k

A graph reasoning benchmark that evaluates language models' ability to find parent nodes in graphs with context length over 128k tokens, testing…

Modality
text
Categories
long_context, reasoning, spatial_reasoning
Openness
unknown
Source
llm_stats
Reported scores
7

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.6Anthropic0.9542026-02-05evidence
Claude Opus 4.8Anthropic0.8332026-05-28evidence
GPT-4.1OpenAI0.252025-04-14evidence
GPT-4.1 miniOpenAI0.112025-04-14evidence
GPT-4.1 nanoOpenAI0.0562025-04-14evidence
GPT-5.4OpenAI0.3242026-03-05evidence
GPT-5.5OpenAI0.5852026-04-23evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.