Benchmark

APEX-Agents

APEX-Agents is a benchmark evaluating AI agents on long horizon professional tasks that require sustained reasoning, planning, and execution across…

Modality
text
Categories
reasoning, agents
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 3.1 ProGoogle0.3352026-02-19evidence
Grok 4.6xAI0.5752026-08-12evidence
Hy3Tencent0.2562026-07-06evidence
Kimi K2.6Moonshot AI0.2792026-04-20evidence
Kimi K3Moonshot AI0.3762026-07-16evidence
MiniMax M3MiniMax0.2772026-06-01evidence
Seed 2.1 ProByteDance0.3382026-06-24evidence
Seed 2.1 TurboByteDance0.2922026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.