Benchmark

SABER Benchmark

What is SABER?

SABER evaluates operational safety of LLM coding agents in stateful project workspaces. Agents perform realistic tasks, and safety is scored from the final…

Released
2026-05-31
Evaluates
Agents, Agents & Tool Use, General AI, Safety
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

SABER paper and code

No reported scores are on record for this benchmark yet.

Which sources cite SABER?

Related Agents, Agents & Tool Use, General AI, Safety benchmarks