Benchmark

MCP-Universe

MCP-Universe evaluates LLMs on complex multi-step agentic tasks using Model Context Protocol (MCP) tools across diverse interactive environments, testing…

Modality
text
Categories
agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3.2DeepSeek0.4592025-12-01evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.