Benchmark
SCALE-QA Benchmark
What is SCALE-QA?
SCALE-QA evaluates interleaved conversational memory using 3,000 audited multiple-choice questions across 10 domains, where correct answers depend on causally…
- Released
- 2026-08-26
- Evaluates
- General AI, Language & Knowledge, Long Context
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-reviewed
- Reported scores
- 0
Source provenance
- Original evidence https://arxiv.org/abs/2608.25655
SCALE-QA paper
No reported scores are on record for this benchmark yet.