Benchmark
BenchBench-Protocol
What is BenchBench-Protocol?
Evaluates LLMs on 149 protocol-modification tasks reconstructed from real changes scientists made to published wet-lab protocols, using weighted rubric scoring.
- Released
- 2026-08-24
- Evaluates
- General AI, Language & Knowledge, Reasoning
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- rule-auto-admitted
- Reported scores
- 0
Source provenance
- Original evidence https://arxiv.org/abs/2608.23898
BenchBench-Protocol paper
No reported scores are on record for this benchmark yet.