Benchmark

Graphwalks parents <128k

A graph reasoning benchmark that evaluates language models' ability to find parent nodes in graphs with context length under 128k tokens, requiring…

Modality
text
Categories
reasoning, spatial_reasoning
Openness
unknown
Source
llm_stats
Reported scores
11

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-4.1OpenAI0.582025-04-14evidence
GPT-4.1 miniOpenAI0.6052025-04-14evidence
GPT-4.1 nanoOpenAI0.0942025-04-14evidence
GPT-4.5OpenAI0.7262025-02-27evidence
GPT-4oOpenAI0.3542024-08-06evidence
GPT-5OpenAI0.7332025-08-07evidence
GPT-5.2OpenAI0.892025-12-11evidence
GPT-5.4OpenAI0.8982026-03-05evidence
GPT-5.4 miniOpenAI0.7152026-03-17evidence
GPT-5.4 nanoOpenAI0.5082026-03-17evidence
o3-miniOpenAI0.5832025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.