Benchmark

Tau3 Banking

τ³-Bench banking domain evaluates agentic models on multi-turn, tool-using customer-support scenarios in a simulated retail banking environment.

Modality
text
Categories
reasoning, agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
7

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Grok 4.5xAI0.332026-07-16evidence
Inkling-SmallThinking Machines Lab0.1552026-07-30evidence
LFM2.5-2.6BLiquid AI0.05672026-08-04evidence
Mistral Medium 3.5Mistral0.1342026-04-29evidence
Muse Glimmer-30BMeta0.2352026-08-10evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.09282026-08-11evidence
Solar Pro 4Upstage0.232026-08-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.