Benchmark

OpenSkillRisk Benchmark

What is OpenSkillRisk?

OpenSkillRisk evaluates LLM-based agents on their ability to recognize and avoid safety risks when using third-party skills, with 263 risky skills in seven…

Released
2026-07-22
Evaluates
Agents, General AI, Safety, Safety & Trustworthiness
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

OpenSkillRisk paper

No reported scores are on record for this benchmark yet.

Which sources cite OpenSkillRisk?

Related Agents, General AI, Safety, Safety & Trustworthiness benchmarks