Benchmark

Humanity's Last Exam

Tool access and test-time compute budget move this score more than model capability does, so two reported figures are rarely comparable. Cards…

Released
2025-01-23
Categories
reasoning
Openness
unknown
Source
model_reports
Reported scores
27

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude 4 OpusAnthropic10.72025-06-17evidence
Claude 4 SonnetAnthropic7.82025-06-17evidence
DeepSeek-V4-FlashDeepSeek34.82026-06-25evidence
DeepSeek-V4-ProDeepSeek7.72026-04-22evidence
DeepSeek-V4-ProDeepSeek34.52026-04-22evidence
DeepSeek-V4-ProDeepSeek44.72026-04-22evidence
DeepSeek-V4-ProDeepSeek37.72026-04-22evidence
DeepSeek-V4-ProDeepSeek48.22026-04-22evidence
DeepSeek-V4-Pro-0813DeepSeek60.02026-08-13evidence
DeepSeek-V4-Pro-0813DeepSeek42.72026-08-13evidence
Gemini 2.5 ProGoogle21.62025-06-17evidence
Gemini 3.1 ProGoogle44.42026-02-19evidence
Gemini 3.1 ProGoogle51.42026-02-19evidence
Gemma 4 (31B)Google19.52026-03-11evidence
Gemma 4 (31B)Google26.52026-03-11evidence
GLM-5Z.ai30.52026-02-11evidence
GLM-5.1Z.ai31.02026-04-03evidence
GLM-5.2Z.ai40.52026-06-16evidence
GLM-5.3-FlashZ.ai55.32026-08-27evidence
Hy4 previewTencent55.42026-08-28evidence
Hy4 previewTencent43.42026-08-28evidence
Kimi K3Moonshot AI56.02026-06-13evidence
Kimi K3Moonshot AI43.52026-06-13evidence
o3 highOpenAI20.32025-06-17evidence
Qwen3.5-397B-A17BQwen28.72026-02-16evidence
Qwen3.5-397B-A17BQwen48.32026-02-16evidence
Qwen3.5-397B-A17BQwen37.62026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.