Benchmark

GDPval

Graded by expert human comparison against real deliverables, and the "-AA" variants are run by Artificial Analysis rather than the vendor, so an Elo here…

Released
2025-09-25
Categories
professional
Openness
unknown
Source
model_reports
Reported scores
6

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude Opus 5Anthropic1861.02026-07-24evidence
DeepSeek-V4-FlashDeepSeek1395.02026-06-25evidence
DeepSeek-V4-ProDeepSeek1554.02026-04-22evidence
GLM-5.3-FlashZ.ai1773.02026-08-27evidence
Hy4 previewTencent1678.02026-08-28evidence
Kimi K3Moonshot AI1686.02026-06-13evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.