Benchmark

GPQA Diamond

198 questions in the Diamond split, so run-to-run variance is wide and a single reported number hides it. Approaching saturation at the 2026 frontier…

Released
2023-11-20
Categories
science
Openness
unknown
Source
model_reports
Reported scores
21

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude 4 OpusAnthropic79.62025-06-17evidence
Claude 4 SonnetAnthropic75.42025-06-17evidence
Claude Opus 4Anthropic79.62025-05-22evidence
Claude Sonnet 4Anthropic75.42025-05-22evidence
DeepSeek-R1DeepSeek71.52025-01-22evidence
DeepSeek-V3DeepSeek59.12024-12-27evidence
DeepSeek-V4-FlashDeepSeek88.12026-06-25evidence
DeepSeek-V4-ProDeepSeek72.92026-04-22evidence
DeepSeek-V4-ProDeepSeek89.12026-04-22evidence
DeepSeek-V4-ProDeepSeek90.12026-04-22evidence
Gemini 2.5 ProGoogle86.42025-06-17evidence
Gemini 3.1 ProGoogle94.32026-02-19evidence
Gemma 4 (31B)Google84.32026-03-11evidence
GLM-5Z.ai86.02026-02-11evidence
GLM-5.1Z.ai86.22026-04-03evidence
GLM-5.2Z.ai91.22026-06-16evidence
Grok 3 Beta Extended ThinkingxAI80.22025-06-17evidence
Hy4 previewTencent92.32026-08-28evidence
Kimi K3Moonshot AI93.52026-06-13evidence
o3 highOpenAI83.32025-06-17evidence
Qwen3-235B-A22B (Thinking)Qwen71.12025-05-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.