Benchmark

CyberGym

CyberGym is a benchmark for evaluating AI agents on cybersecurity tasks, testing their ability to identify vulnerabilities, perform security analysis, and…

Modality
text
Categories
safety, agents, code
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Mythos PreviewAnthropic0.8312026-04-07evidence
Claude Opus 4.6Anthropic0.7382026-02-05evidence
Claude Opus 4.7Anthropic0.7312026-04-16evidence
Claude Opus 4.8Anthropic0.7882026-05-28evidence
DeepSeek-V4-Flash-0731DeepSeek0.7672026-07-31evidence
DeepSeek-V4-Pro-0813DeepSeek0.8332026-08-13evidence
Gemini 3.5 Flash CyberGoogle0.8322026-07-21evidence
GLM-5.1Z.ai0.6872026-04-07evidence
GLM-5.3Z.ai0.8452026-08-14evidence
GPT-5.5OpenAI0.8182026-04-23evidence
Kimi K2.5Moonshot AI0.4132026-01-27evidence
Seed 2.1 ProByteDance0.7022026-06-24evidence
Seed 2.1 TurboByteDance0.672026-06-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.