Benchmark

t2-bench

t2-bench is a benchmark for evaluating agentic tool use capabilities, measuring how well models can select, sequence, and utilize tools to solve complex…

Modality
text
Categories
reasoning, agents, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
23

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3.2DeepSeek0.8032025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.8022025-12-01evidence
DeepSeek-V3.2-SpecialeDeepSeek0.8032025-12-01evidence
DiffusionGemma 26B-A4BGoogle0.5622026-06-10evidence
Gemini 3 FlashGoogle0.9022025-12-17evidence
Gemini 3 ProGoogle0.8542025-11-18evidence
Gemini 3.1 ProGoogle0.9932026-02-19evidence
Gemma 4 26B-A4BGoogle0.8552026-04-02evidence
Gemma 4 31BGoogle0.8642026-04-02evidence
Gemma 4 E2BGoogle0.2942026-04-02evidence
Gemma 4 E4BGoogle0.5752026-04-02evidence
GLM-5Z.ai0.8972026-02-11evidence
GPT OSS 120B HighOpenAI0.6392025-08-05evidence
K-EXAONE-236B-A23BLG AI Research0.7322025-12-31evidence
Qwen3 MaxQwen0.7482025-12-15evidence
Qwen3.5-0.8BQwen0.1162026-03-02evidence
Qwen3.5-122B-A10BQwen0.7952026-02-24evidence
Qwen3.5-27BQwen0.792026-02-24evidence
Qwen3.5-2BQwen0.4882026-03-02evidence
Qwen3.5-35B-A3BQwen0.8122026-02-24evidence
Qwen3.5-397B-A17BQwen0.8672026-02-16evidence
Qwen3.5-4BQwen0.7992026-03-02evidence
Qwen3.5-9BQwen0.7912026-03-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.