Benchmark

Internal Research Debugging Evaluation

The Internal Research Debugging Evaluation measures whether models can debug 41 real bugs from internal OpenAI research experiments (plus…

Modality
text
Categories
reasoning, agents, code
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5.6 LunaOpenAI0.5082026-07-09evidence
GPT-5.6 SolOpenAI0.6832026-07-09evidence
GPT-5.6 TerraOpenAI0.6782026-07-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.