Benchmark

Long-Horizon-Terminal-Bench

What is Long-Horizon-Terminal-Bench?

Evaluates long-horizon terminal tasks in a containerized environment with hidden verifiers and dense reward grading across 46 tasks and nine categories.

Released
2026-07-09
Evaluates
Agents & Tool Use, Code, General AI, Multimodal, Software & AI Compute
Openness
unknown
Importer
Claire Radar
Review status
source-reviewed
Reported scores
0

Source provenance

Long-Horizon-Terminal-Bench paper, code and dataset

No reported scores are on record for this benchmark yet.

Which sources cite Long-Horizon-Terminal-Bench?

Related Agents & Tool Use, Code, General AI, Multimodal, Software & AI Compute benchmarks