Benchmark

BTS-AgentBench

What is BTS-AgentBench?

Evaluates multi-turn agent performance on 532 telemetry-derived tasks across train/dev/test splits, with additional XAI4HEAT episodes, using deterministic…

Released
2026-08-27
Evaluates
Agents, General AI, Language & Knowledge
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

BTS-AgentBench paper and code

No reported scores are on record for this benchmark yet.

Which sources cite BTS-AgentBench?

Related Agents, General AI, Language & Knowledge benchmarks