Benchmark

Humanity's Last Exam (with tools, text-only)

Text-only Humanity's Last Exam variant evaluated with tool use enabled.

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-Pro-0813DeepSeek0.62026-08-13evidence
Hy3Tencent0.5322026-07-06evidence
Qwen3.8 MaxQwen0.5622026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.