Benchmark

Tau2 Retail

τ²-bench retail domain evaluates conversational AI agents in customer service scenarios within a dual-control environment where both agent and user can…

Modality
text
Categories
reasoning, communication, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
26

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Haiku 4.5Anthropic0.8322025-10-15evidence
Claude Opus 4.5Anthropic0.8892025-11-24evidence
Claude Opus 4.6Anthropic0.9192026-02-05evidence
Claude Sonnet 4.6Anthropic0.9172026-02-17evidence
GPT-4oOpenAI0.6342024-08-06evidence
GPT-5OpenAI0.8112025-08-07evidence
GPT-5.1OpenAI0.7792025-11-13evidence
GPT-5.1 InstantOpenAI0.7792025-11-12evidence
GPT-5.1 ThinkingOpenAI0.7792025-11-12evidence
GPT-5.2OpenAI0.822025-12-11evidence
Kimi K2 InstructMoonshot AI0.7062025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.7062025-09-05evidence
LongCat-Flash-ChatMeituan0.71272025-08-29evidence
LongCat-Flash-LiteMeituan0.7312026-02-05evidence
LongCat-Flash-ThinkingMeituan0.7152025-09-22evidence
LongCat-Flash-Thinking-2601Meituan0.8862026-01-14evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.5692025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.62832026-03-11evidence
Nova 2 LiteAmazon0.7652025-12-02evidence
Nova 2 OmniAmazon0.7832025-12-02evidence
Nova 2 ProAmazon0.7772025-12-02evidence
o3OpenAI0.8022025-04-16evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.7132025-07-22evidence
Qwen3-235B-A22B-Thinking-2507Qwen0.7192025-07-25evidence
Qwen3-Next-80B-A3B-InstructQwen0.5732025-09-10evidence
Qwen3-Next-80B-A3B-ThinkingQwen0.6782025-09-10evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.