Benchmark

BrowseComp Long Context 128k

A challenging benchmark for evaluating web browsing agents' ability to persistently navigate the internet and find hard-to-locate, entangled information…

Modality
text
Categories
reasoning, search
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5OpenAI0.92025-08-07evidence
GPT-5.1OpenAI0.92025-11-13evidence
GPT-5.1 InstantOpenAI0.92025-11-12evidence
GPT-5.1 ThinkingOpenAI0.92025-11-12evidence
GPT-5.2OpenAI0.922025-12-11evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.