Benchmark

AgentDrift Benchmark

What is AgentDrift?

Step-labeled corpus of 12,536 synthetic LLM agent tool-call trajectories in five domains, labeled per step for benign, injection point, hijacked, or failed…

Released
2026-09-07
Evaluates
Agents, General AI, Language & Knowledge
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

AgentDrift paper and code

No reported scores are on record for this benchmark yet.

Which sources cite AgentDrift?

Related Agents, General AI, Language & Knowledge benchmarks