Benchmark

Graphwalks BFS <128k

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length under…

Modality
text
Categories
reasoning, spatial_reasoning
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-4.1OpenAI0.6172025-04-14evidence
GPT-4.1 miniOpenAI0.6172025-04-14evidence
GPT-4.1 nanoOpenAI0.252025-04-14evidence
GPT-4.5OpenAI0.7232025-02-27evidence
GPT-4oOpenAI0.4172024-08-06evidence
GPT-5OpenAI0.7832025-08-07evidence
GPT-5.2OpenAI0.942025-12-11evidence
GPT-5.4OpenAI0.932026-03-05evidence
GPT-5.4 miniOpenAI0.7632026-03-17evidence
GPT-5.4 nanoOpenAI0.7342026-03-17evidence
o3-miniOpenAI0.512025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.