Benchmark

TAU3-Bench

TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models…

Modality
text
Categories
reasoning, agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
5

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GLM-5.1Z.ai0.7062026-04-07evidence
MiMo-V2.5-ProXiaomi0.7292026-04-27evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.2262026-06-04evidence
Qwen3.6 PlusQwen0.7072026-04-02evidence
Qwen3.6-35B-A3BQwen0.6722026-04-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.