Benchmark

Tau2 Airline

TAU2 airline domain benchmark for evaluating conversational agents in dual-control environments where both AI agents and users interact with tools in…

Modality
text
Categories
reasoning, communication, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
23

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Haiku 4.5Anthropic0.6362025-10-15evidence
GPT-4oOpenAI0.4552024-08-06evidence
GPT-5OpenAI0.6262025-08-07evidence
GPT-5.1OpenAI0.672025-11-13evidence
GPT-5.1 InstantOpenAI0.672025-11-12evidence
GPT-5.1 ThinkingOpenAI0.672025-11-12evidence
Kimi K2 InstructMoonshot AI0.5652025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.5652025-09-05evidence
LongCat-Flash-ChatMeituan0.582025-08-29evidence
LongCat-Flash-LiteMeituan0.582026-02-05evidence
LongCat-Flash-ThinkingMeituan0.6752025-09-22evidence
LongCat-Flash-Thinking-2601Meituan0.7652026-01-14evidence
Mercury 2Inception0.532026-02-24evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.482025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.56252026-03-11evidence
Nova 2 LiteAmazon0.6482025-12-02evidence
Nova 2 OmniAmazon0.6882025-12-02evidence
Nova 2 ProAmazon0.6522025-12-02evidence
o3OpenAI0.6482025-04-16evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.442025-07-22evidence
Qwen3-235B-A22B-Thinking-2507Qwen0.582025-07-25evidence
Qwen3-Next-80B-A3B-InstructQwen0.4552025-09-10evidence
Qwen3-Next-80B-A3B-ThinkingQwen0.6052025-09-10evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.