Benchmark

context-flip evaluation

What is context-flip evaluation?

Context-flip evaluation tests whether aligned language models adjust their actions when a situational change reverses which action is safe.

Released
2026-05-27
Evaluates
General AI, Safety, Safety & Trustworthiness
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

context-flip evaluation paper

No reported scores are on record for this benchmark yet.

Which sources cite context-flip evaluation?

Related General AI, Safety, Safety & Trustworthiness benchmarks