Benchmark

SkillTV-Bench

What is SkillTV-Bench?

Evaluates whether a judge can verify long-horizon agent executions using original instructions, normalized trajectories, task-time skills, and inspectable…

Released
2026-08-06
Evaluates
Agents, General AI, Language & Knowledge
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

SkillTV-Bench paper and code

No reported scores are on record for this benchmark yet.

Which sources cite SkillTV-Bench?

Related Agents, General AI, Language & Knowledge benchmarks