Benchmark

Failure-Transparent Agents Benchmark

What is Failure-Transparent Agents?

Measures whether LLM agents admit tool failures instead of claiming success across 100 tasks covering web, attachment, execution, permission, and stale-data…

Released
2026-09-10
Evaluates
Language & Knowledge, Software & AI Compute, cs.AI
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

Failure-Transparent Agents code

No reported scores are on record for this benchmark yet.

Which sources cite Failure-Transparent Agents?

Related Language & Knowledge, Software & AI Compute, cs.AI benchmarks