Benchmark
TAU3-Bench
TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models…
- Modality
- text
- Categories
- reasoning, agents, tool_calling
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 5
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| GLM-5.1 | Z.ai | 0.706 | 2026-04-07 | evidence |
| MiMo-V2.5-Pro | Xiaomi | 0.729 | 2026-04-27 | evidence |
| Nemotron 3 Ultra (550B A55B) | NVIDIA | 0.226 | 2026-06-04 | evidence |
| Qwen3.6 Plus | Qwen | 0.707 | 2026-04-02 | evidence |
| Qwen3.6-35B-A3B | Qwen | 0.672 | 2026-04-16 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.