Benchmark
KC-Bench
What is KC-Bench?
A multi-turn benchmark with 238 tasks using a user simulator, stateful tools, and environment assertions to test LLM agents' handling of knowledge conflicts…
- Released
- 2026-09-03
- Evaluates
- Agents & Tool Use, Factuality, General AI, cs.AI
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-name-audit-deferred
- Reported scores
- 0
Source provenance
- Original evidence https://arxiv.org/abs/2609.03588
KC-Bench paper
No reported scores are on record for this benchmark yet.