Benchmark

Terminal-Bench 2.1

Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It…

Modality
text
Categories
reasoning, agents, code, tool_calling
Openness
unknown
Source
llm_stats
Reported scores
28

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.8432026-06-09evidence
DeepSeek-V4-Flash-0731DeepSeek0.8272026-07-31evidence
DeepSeek-V4-Pro-0813DeepSeek0.8792026-08-13evidence
Gemini 3.5 Flash-LiteGoogle0.542026-07-21evidence
Gemini 3.6 FlashGoogle0.782026-07-21evidence
Gemini 3.7 FlashGoogle0.8582026-08-13evidence
GLM-5.2Z.ai0.8272026-06-16evidence
GLM-5.3Z.ai0.8822026-08-14evidence
GPT-5.6 LunaOpenAI0.8472026-07-09evidence
GPT-5.6 SolOpenAI0.8882026-07-09evidence
GPT-5.6 TerraOpenAI0.8742026-07-09evidence
Grok 4.5xAI0.8332026-07-16evidence
Hy3Tencent0.7172026-07-06evidence
Inkling-SmallThinking Machines Lab0.6472026-07-30evidence
Kimi K3Moonshot AI0.8832026-07-16evidence
Laguna S 2.1Poolside0.7022026-07-21evidence
MAI-Code-1.1-FlashMicrosoft0.6292026-08-11evidence
MiniMax M3MiniMax0.662026-06-01evidence
Muse Glimmer-30BMeta0.5172026-08-10evidence
Muse Spark 1.1Meta0.82026-07-09evidence
Muse Spark 1.2Meta0.8292026-08-05evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.5642026-06-04evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.24582026-08-11evidence
Qwen3.8 MaxQwen0.8662026-08-02evidence
Qwen3.8-27BQwen0.732026-08-14evidence
Seed 2.1 ProByteDance0.712026-06-24evidence
Seed 2.1 TurboByteDance0.6762026-06-24evidence
Solar Pro 4Upstage0.572026-08-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.