Benchmark

REDAgentBench

What is REDAgentBench?

Evaluates LLM agent safety through executable red-teaming, adversarial case generation, and verification of harmful effects across 1,661 cases and five service…

Released
2026-08-11
Evaluates
Agents, Agents & Tool Use, Cybersecurity
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

REDAgentBench paper

No reported scores are on record for this benchmark yet.

Which sources cite REDAgentBench?

Related Agents, Agents & Tool Use, Cybersecurity benchmarks