Benchmark

Tau2 Telecom

τ²-Bench telecom domain evaluates conversational agents in a dual-control environment modeled as a Dec-POMDP, where both agent and user use tools in…

Modality
text
Categories
reasoning, communication, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
35

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Haiku 4.5Anthropic0.832025-10-15evidence
Claude Opus 4.5Anthropic0.9822025-11-24evidence
Claude Opus 4.6Anthropic0.9932026-02-05evidence
Claude Sonnet 4.6Anthropic0.9792026-02-17evidence
Command A+Cohere0.852026-05-20evidence
GPT-4oOpenAI0.2352024-08-06evidence
GPT-5OpenAI0.9672025-08-07evidence
GPT-5.1OpenAI0.9562025-11-13evidence
GPT-5.1 InstantOpenAI0.9562025-11-12evidence
GPT-5.1 ThinkingOpenAI0.9562025-11-12evidence
GPT-5.2OpenAI0.9872025-12-11evidence
GPT-5.4OpenAI0.9892026-03-05evidence
GPT-5.4 miniOpenAI0.9342026-03-17evidence
GPT-5.4 nanoOpenAI0.9252026-03-17evidence
GPT-5.5OpenAI0.982026-04-23evidence
Kimi K2 InstructMoonshot AI0.6582025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.6582025-09-05evidence
LongCat-Flash-ChatMeituan0.73682025-08-29evidence
LongCat-Flash-LiteMeituan0.7282026-02-05evidence
LongCat-Flash-ThinkingMeituan0.8312025-09-22evidence
LongCat-Flash-Thinking-2601Meituan0.9932026-01-14evidence
MAI-Code-1-FlashMicrosoft0.7172026-06-02evidence
MiMo-V2-ProXiaomi0.9682026-03-18evidence
MiniMax M2MiniMax0.872025-10-27evidence
MiniMax M2.1MiniMax0.872025-12-23evidence
Muse SparkMeta0.9152026-04-08evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.4222025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.64362026-03-11evidence
Nova 2 LiteAmazon0.762025-12-02evidence
Nova 2 OmniAmazon0.82025-12-02evidence
Nova 2 ProAmazon0.9272025-12-02evidence
o3OpenAI0.5822025-04-16evidence
Qwen3-235B-A22B-Thinking-2507Qwen0.4562025-07-25evidence
Qwen3-Next-80B-A3B-InstructQwen0.1322025-09-10evidence
Qwen3-Next-80B-A3B-ThinkingQwen0.4392025-09-10evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.