Benchmark

Graphwalks BFS >128k

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over…

Modality
text
Categories
long_context, reasoning, spatial_reasoning
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Mythos PreviewAnthropic0.82026-04-07evidence
Claude Opus 4.6Anthropic0.6152026-02-05evidence
Claude Opus 4.8Anthropic0.6812026-05-28evidence
GPT-4.1OpenAI0.192025-04-14evidence
GPT-4.1 miniOpenAI0.152025-04-14evidence
GPT-4.1 nanoOpenAI0.0292025-04-14evidence
GPT-5.4OpenAI0.2142026-03-05evidence
GPT-5.5OpenAI0.4542026-04-23evidence
GPT-5.6 LunaOpenAI0.8132026-07-09evidence
GPT-5.6 SolOpenAI0.9072026-07-09evidence
GPT-5.6 TerraOpenAI0.7692026-07-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.