Benchmark

SchemeArena Benchmark

What is SchemeArena?

Evaluates scheming in LLM agents across 400 factorized scenarios, with a monitor that scores covert misaligned behavior along five rubric dimensions.

Released
2026-09-08
Evaluates
General AI, Safety, Safety & Trustworthiness
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

SchemeArena paper and code

No reported scores are on record for this benchmark yet.

Which sources cite SchemeArena?

Related General AI, Safety, Safety & Trustworthiness benchmarks