Benchmark

FrontierChallenge Benchmark

What is FrontierChallenge?

Pass-rate results are model-scaffold measurements over the 97 released tasks with one trajectory per system-task pair, so a reported percentage is not a…

Released
2026-08-25
Evaluates
scientific agent
Openness
unknown
Catalog source
Model reports
Reported scores
13

FrontierChallenge links

FrontierChallenge results and reported scores

Model reports

ModelOrganizationReported valueReportedEvidence
Apodex 1.1ApodexAI10.32026-08-25evidence
Apodex 1.1ApodexAI12.42026-08-25evidence
Claude Opus 5Anthropic17.52026-08-25evidence
DeepSeek V4 Flash-0731DeepSeek12.42026-08-25evidence
DeepSeek V4 Pro-0813DeepSeek13.42026-08-25evidence
Gemini 3.7 FlashGoogle10.32026-08-25evidence
GLM-5.2Z.ai3.12026-08-25evidence
GPT-5.6 SolOpenAI20.62026-08-25evidence
GPT-5.6 Terra (max)OpenAI15.52026-08-25evidence
Grok 4.6xAI20.62026-08-25evidence
Kimi K3Moonshot AI17.52026-08-25evidence
Qwen 3.8 MaxQwen15.52026-08-25evidence
Qwen3.5-397B-A17BQwen4.12026-08-25evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.

Which sources cite FrontierChallenge?