Benchmark

MCP Atlas

MCP Atlas is a benchmark for evaluating AI models on scaled tool use capabilities, measuring how well models can coordinate and utilize multiple tools…

Modality
text
Categories
reasoning, agents, code, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
33

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Opus 4.5Anthropic0.6232025-11-24evidence
Claude Opus 4.6Anthropic0.6272026-02-05evidence
Claude Opus 4.7Anthropic0.7732026-04-16evidence
Claude Opus 4.8Anthropic0.8222026-05-28evidence
Claude Sonnet 4.6Anthropic0.6132026-02-17evidence
DeepSeek-V4-Flash-0423DeepSeek0.6742026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.692026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.7362026-04-23evidence
Gemini 3 FlashGoogle0.5742025-12-17evidence
Gemini 3.1 ProGoogle0.6922026-02-19evidence
Gemini 3.5 FlashGoogle0.8362026-05-19evidence
GLM-5Z.ai0.6782026-02-11evidence
GLM-5.1Z.ai0.7182026-04-07evidence
GLM-5.2Z.ai0.7682026-06-16evidence
GPT-5.2OpenAI0.6062025-12-11evidence
GPT-5.4OpenAI0.6722026-03-05evidence
GPT-5.4 miniOpenAI0.5772026-03-17evidence
GPT-5.4 nanoOpenAI0.5612026-03-17evidence
GPT-5.5OpenAI0.7532026-04-23evidence
Hy3Tencent0.7912026-07-06evidence
Inkling-SmallThinking Machines Lab0.7962026-07-30evidence
Kimi K2.7 CodeMoonshot AI0.762026-06-12evidence
Kimi K3Moonshot AI0.8422026-07-16evidence
MiniMax M3MiniMax0.7422026-06-01evidence
Muse Glimmer-30BMeta0.7552026-08-10evidence
Muse Spark 1.1Meta0.8812026-07-09evidence
Nova 2 LiteAmazon0.2462025-12-02evidence
Qwen3.6 PlusQwen0.7412026-04-02evidence
Qwen3.6-35B-A3BQwen0.6282026-04-16evidence
Qwen3.7 MaxQwen0.7642026-05-19evidence
Qwen3.7-PlusQwen0.7322026-05-31evidence
Seed 2.1 ProByteDance0.8382026-06-24evidence
Seed 2.1 TurboByteDance0.8032026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.