Benchmark
context-flip evaluation
What is context-flip evaluation?
Context-flip evaluation tests whether aligned language models adjust their actions when a situational change reverses which action is safe.
- Released
- 2026-05-27
- Evaluates
- General AI, Safety, Safety & Trustworthiness
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-reviewed
- Reported scores
- 0
Source provenance
- Original evidence https://arxiv.org/abs/2605.27851
context-flip evaluation paper
No reported scores are on record for this benchmark yet.