Benchmark
ClawProBench
What is ClawProBench?
Evaluates agent configurations on 102 live-runtime and 68 frozen holdout scenarios, scoring execution traces with a safety-gated formula covering correctness…
- Released
- 2026-08-23
- Evaluates
- Agents, General AI, Language & Knowledge
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- source-reviewed
- Reported scores
- 0
Source provenance
- Original evidence http://arxiv.org/abs/2608.22510v1
ClawProBench paper and dataset
No reported scores are on record for this benchmark yet.