Benchmark

ToolRobustBench

What is ToolRobustBench?

Evaluates tool-calling agents under perturbations across the tool-use pipeline, attributing failures to selection, grounding, argument binding, and feedback…

Released
2026-08-23
Evaluates
Agents, Code & Software, General AI
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

ToolRobustBench paper

No reported scores are on record for this benchmark yet.

Which sources cite ToolRobustBench?

Related Agents, Code & Software, General AI benchmarks