Benchmark

BAITBENCH

What is BAITBENCH?

Evaluates whether LLM agents exploit optional shortcuts in three synthetic tabular ML tasks to inflate public test scores while failing hidden tests.

Released
2026-08-31
Evaluates
Agents, Cybersecurity, Language & Knowledge
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

BAITBENCH paper

No reported scores are on record for this benchmark yet.

Which sources cite BAITBENCH?

Related Agents, Cybersecurity, Language & Knowledge benchmarks