Benchmark

MCP-Mark

MCP-Mark evaluates LLMs on their ability to use Model Context Protocol (MCP) tools effectively, testing tool discovery, selection, invocation, and result…

Modality
text
Categories
agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3.2DeepSeek0.382025-12-01evidence
Kimi K2.6Moonshot AI0.5592026-04-20evidence
Kimi K2.7 CodeMoonshot AI0.8112026-06-12evidence
Qwen3.5-397B-A17BQwen0.4612026-02-16evidence
Qwen3.6 PlusQwen0.4822026-04-02evidence
Qwen3.6-35B-A3BQwen0.372026-04-16evidence
Qwen3.7 MaxQwen0.6082026-05-19evidence
Qwen3.7-PlusQwen0.5872026-05-31evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.