Benchmark

Terminal-Bench Hard

Terminal-Bench Hard is a harder terminal-agent benchmark variant evaluated with the Terminus-2 harness in Cohere's Command A+ and North Mini Code releases.

Modality
text
Categories
reasoning, agents, code, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Command A+Cohere0.252026-05-20evidence
North Mini Code 1.0Cohere0.3112026-06-09evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.