Benchmark

AutomationBench

AutomationBench is a tool-use benchmark that evaluates AI agents on automating real-world workflows, testing their ability to orchestrate tools and…

Modality
text
Categories
reasoning, agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
12

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.1742026-06-09evidence
Claude Opus 5Anthropic0.262026-07-24evidence
Claude Sonnet 5Anthropic0.1352026-06-30evidence
DeepSeek-V4-Flash-0731DeepSeek0.2512026-07-31evidence
DeepSeek-V4-Pro-0813DeepSeek0.3182026-08-13evidence
Gemini 3.7 FlashGoogle0.3042026-08-13evidence
GLM-5.3Z.ai0.4822026-08-14evidence
GPT-5.6 LunaOpenAI0.1492026-07-09evidence
GPT-5.6 SolOpenAI0.1812026-07-09evidence
GPT-5.6 TerraOpenAI0.1522026-07-09evidence
Kimi K3Moonshot AI0.3082026-07-16evidence
Qwen3.8 MaxQwen0.2732026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.