Benchmark

APEX-Agents

Expert-authored professional tasks with rubric grading; graders are LLMs.

Released
2026-01-20
Categories
agent
Openness
unknown
Source
model_reports
Reported scores
5

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-FlashDeepSeek33.02026-06-25evidence
DeepSeek-V4-ProDeepSeek38.32026-04-22evidence
Gemini 3.1 ProGoogle33.52026-02-19evidence
Hy4 previewTencent37.12026-08-28evidence
Kimi K3Moonshot AI41.02026-06-13evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.