Benchmark

AgentRelBench

What is AgentRelBench?

AgentRelBench measures repeated agent runs in a fixed task suite, computing severity-weighted damage from database state diffs with no LLM judging.

Released
2026-08-15
Evaluates
Agents, Agents & Tool Use, General AI
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

AgentRelBench paper and code

No reported scores are on record for this benchmark yet.

Which sources cite AgentRelBench?

Related Agents, Agents & Tool Use, General AI benchmarks