Benchmark

EdgeBench

What is EdgeBench?

EdgeBench evaluates autonomous agents on 134 real-world tasks across scientific discovery, software engineering, optimization, knowledge work, formal…

Released
2026-07-02
Evaluates
Agents, Agents & Tool Use, Environment Learning, Environment learning, General AI, Learning from feedback, Long-horizon task execution, Mathematics & Formal Science, Scaling Laws, Scientific Research & AI for Science, Software & AI Compute
Openness
unknown
Importer
Claire Radar
Review status
source-reviewed
Reported scores
0

Source provenance

EdgeBench paper, code and dataset

No reported scores are on record for this benchmark yet.

Which sources cite EdgeBench?

Related Agents, Agents & Tool Use, Environment Learning, Environment learning, General AI, Learning from feedback, Long-horizon task execution, Mathematics & Formal Science, Scaling Laws, Scientific Research & AI for Science, Software & AI Compute benchmarks