Benchmark
GDPval
Graded by expert human comparison against real deliverables, and the "-AA" variants are run by Artificial Analysis rather than the vendor, so an Elo here…
- Released
- 2025-09-25
- Categories
- professional
- Openness
- unknown
- Source
- model_reports
- Reported scores
- 6
Reported scores
model_reports
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 1861.0 | 2026-07-24 | evidence |
| DeepSeek-V4-Flash | DeepSeek | 1395.0 | 2026-06-25 | evidence |
| DeepSeek-V4-Pro | DeepSeek | 1554.0 | 2026-04-22 | evidence |
| GLM-5.3-Flash | Z.ai | 1773.0 | 2026-08-27 | evidence |
| Hy4 preview | Tencent | 1678.0 | 2026-08-28 | evidence |
| Kimi K3 | Moonshot AI | 1686.0 | 2026-06-13 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.