Benchmark
Failure-Transparent Agents Benchmark
What is Failure-Transparent Agents?
Measures whether LLM agents admit tool failures instead of claiming success across 100 tasks covering web, attachment, execution, permission, and stale-data…
- Released
- 2026-09-10
- Evaluates
- Language & Knowledge, Software & AI Compute, cs.AI
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-reviewed
- Reported scores
- 0
Source provenance
- Original evidence https://github.com/junru-zhu/failure-transparent-agents
Failure-Transparent Agents code
No reported scores are on record for this benchmark yet.