Benchmark

Terminal-Bench

Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks…

Modality
text
Categories
reasoning, agents, code
Openness
unknown
Source
llm_stats
Reported scores
25

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3.7 SonnetAnthropic0.3522025-02-24evidence
Claude Haiku 4.5Anthropic0.412025-10-15evidence
Claude Opus 4Anthropic0.3922025-05-22evidence
Claude Opus 4.1Anthropic0.4332025-08-05evidence
Claude Sonnet 4Anthropic0.3552025-05-22evidence
Claude Sonnet 4.5Anthropic0.52025-09-29evidence
DeepSeek-R1-0528DeepSeek0.0572025-05-28evidence
DeepSeek-V3.1DeepSeek0.3132025-01-10evidence
DeepSeek-V3.2-ExpDeepSeek0.3772025-09-29evidence
GLM-4.5Z.ai0.3752025-07-28evidence
GLM-4.5-AirZ.ai0.32025-07-28evidence
GLM-4.6Z.ai0.4052025-09-30evidence
GLM-4.7Z.ai0.3332025-12-22evidence
Kimi K2 InstructMoonshot AI0.32025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.252025-09-05evidence
Kimi K2-Thinking-0905Moonshot AI0.4712025-09-05evidence
LongCat-Flash-ChatMeituan0.39512025-08-29evidence
LongCat-Flash-LiteMeituan0.33752026-02-05evidence
MiMo-V2-FlashXiaomi0.3052025-12-16evidence
MiniMax M2MiniMax0.4632025-10-27evidence
MiniMax M2.1MiniMax0.4792025-12-23evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.0852025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.25782026-03-11evidence
Nova 2 LiteAmazon0.3252025-12-02evidence
Nova 2 ProAmazon0.4132025-12-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.