Benchmark
FrontierChallenge Benchmark
What is FrontierChallenge?
Pass-rate results are model-scaffold measurements over the 97 released tasks with one trajectory per system-task pair, so a reported percentage is not a…
- Released
- 2026-08-25
- Evaluates
- scientific agent
- Openness
- unknown
- Catalog source
- Model reports
- Reported scores
- 13
FrontierChallenge links
FrontierChallenge results and reported scores
Model reports
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Apodex 1.1 | ApodexAI | 10.3 | 2026-08-25 | evidence |
| Apodex 1.1 | ApodexAI | 12.4 | 2026-08-25 | evidence |
| Claude Opus 5 | Anthropic | 17.5 | 2026-08-25 | evidence |
| DeepSeek V4 Flash-0731 | DeepSeek | 12.4 | 2026-08-25 | evidence |
| DeepSeek V4 Pro-0813 | DeepSeek | 13.4 | 2026-08-25 | evidence |
| Gemini 3.7 Flash | 10.3 | 2026-08-25 | evidence | |
| GLM-5.2 | Z.ai | 3.1 | 2026-08-25 | evidence |
| GPT-5.6 Sol | OpenAI | 20.6 | 2026-08-25 | evidence |
| GPT-5.6 Terra (max) | OpenAI | 15.5 | 2026-08-25 | evidence |
| Grok 4.6 | xAI | 20.6 | 2026-08-25 | evidence |
| Kimi K3 | Moonshot AI | 17.5 | 2026-08-25 | evidence |
| Qwen 3.8 Max | Qwen | 15.5 | 2026-08-25 | evidence |
| Qwen3.5-397B-A17B | Qwen | 4.1 | 2026-08-25 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.
Which sources cite FrontierChallenge?
- FrontierChallenge leaderboard ApodexAI