Benchmark

AFT-Bench

What is AFT-Bench?

Holds task, backend, initial state, injected failure, agent, and language model fixed while varying the tool interface to measure callability versus…

Released
2026-08-23
Evaluates
Agents, Code & Software, General AI
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

AFT-Bench paper

No reported scores are on record for this benchmark yet.

Which sources cite AFT-Bench?

Related Agents, Code & Software, General AI benchmarks