Benchmark

Terminal-Bench

Environment image, timeout and permitted commands change results independently of the model. Version and harness both matter: 2.0 and 2.1 are different…

Released
2025-05-19
Categories
agent
Openness
unknown
Source
model_reports
Reported scores
15

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
Claude Opus 4Anthropic43.22025-05-22evidence
DeepSeek-V4-FlashDeepSeek56.92026-06-25evidence
DeepSeek-V4-ProDeepSeek59.12026-04-22evidence
DeepSeek-V4-ProDeepSeek63.32026-04-22evidence
DeepSeek-V4-ProDeepSeek67.92026-04-22evidence
DeepSeek-V4-Pro-0813DeepSeek87.92026-08-13evidence
Gemini 3.1 ProGoogle68.52026-02-19evidence
GLM-5Z.ai56.22026-02-11evidence
GLM-5.1Z.ai63.52026-04-03evidence
GLM-5.2Z.ai81.02026-06-16evidence
GLM-5.3-FlashZ.ai84.32026-08-27evidence
Grok 4.5xAI83.32026-07-08evidence
Hy4 previewTencent85.42026-08-28evidence
Kimi K3Moonshot AI88.32026-06-13evidence
Qwen3.5-397B-A17BQwen52.52026-02-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.