Benchmark
LLAMIA-Bench
What is LLAMIA-Bench?
Evaluates LLM collaboration with non-language chess engines across six tasks covering imitation, assessment, and explanation, scored on task performance and…
- Released
- 2026-08-31
- Evaluates
- Agents, General AI, Language & Knowledge
- Openness
- unknown
- Importer
- Claire Radar
- Review status
- ai-name-audit-deferred
- Reported scores
- 0
Source provenance
- Original evidence https://arxiv.org/abs/2609.00474
LLAMIA-Bench paper
No reported scores are on record for this benchmark yet.