Benchmark

InterveneBench

What is InterveneBench?

Evaluates human-in-the-loop safety in tool-using agents across 1,000 cases from 250 synthetic tasks with four fixed human replies, covering 24 domains and 20…

Released
2026-09-12
Evaluates
Agents, General AI, Safety, Safety & Trustworthiness
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

InterveneBench code

No reported scores are on record for this benchmark yet.

Which sources cite InterveneBench?

Related Agents, General AI, Safety, Safety & Trustworthiness benchmarks