Benchmark

Toolathlon

Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures…

Modality
text
Categories
reasoning, agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
37

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.8Anthropic0.5992026-05-28evidence
Claude Sonnet 5Anthropic0.5432026-06-30evidence
DeepSeek-V3.2DeepSeek0.3522025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.3522025-12-01evidence
DeepSeek-V3.2-SpecialeDeepSeek0.3522025-12-01evidence
DeepSeek-V4-Flash-0423DeepSeek0.4352026-04-23evidence
DeepSeek-V4-Flash-0731DeepSeek0.7032026-07-31evidence
DeepSeek-V4-Flash-MaxDeepSeek0.4782026-04-23evidence
DeepSeek-V4-Pro-0813DeepSeek0.7412026-08-13evidence
DeepSeek-V4-Pro-MaxDeepSeek0.5182026-04-23evidence
Gemini 3 FlashGoogle0.4942025-12-17evidence
Gemini 3.5 FlashGoogle0.5652026-05-19evidence
GLM-5.1Z.ai0.4072026-04-07evidence
GLM-5.2Z.ai0.4822026-06-16evidence
GLM-5.3Z.ai0.732026-08-14evidence
GPT-5.2OpenAI0.4632025-12-11evidence
GPT-5.4OpenAI0.5462026-03-05evidence
GPT-5.4 miniOpenAI0.4292026-03-17evidence
GPT-5.4 nanoOpenAI0.3552026-03-17evidence
GPT-5.5OpenAI0.5562026-04-23evidence
GPT-5.6 LunaOpenAI0.5342026-07-09evidence
GPT-5.6 SolOpenAI0.582026-07-09evidence
GPT-5.6 TerraOpenAI0.5312026-07-09evidence
Hy3Tencent0.4852026-07-06evidence
Inkling-SmallThinking Machines Lab0.5442026-07-30evidence
Kimi K2.6Moonshot AI0.52026-04-20evidence
Kimi K3Moonshot AI0.7322026-07-16evidence
Laguna S 2.1Poolside0.4972026-07-21evidence
MiniMax M2.1MiniMax0.4352025-12-23evidence
MiniMax M2.7MiniMax0.4632026-03-18evidence
Muse Spark 1.1Meta0.7562026-07-09evidence
Qwen3.5-397B-A17BQwen0.3832026-02-16evidence
Qwen3.6 PlusQwen0.3982026-04-02evidence
Qwen3.6-35B-A3BQwen0.2692026-04-16evidence
Qwen3.8 MaxQwen0.7252026-08-02evidence
Seed 2.1 ProByteDance0.5062026-06-24evidence
Seed 2.1 TurboByteDance0.4912026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.