Benchmark

LongWoF-Bench

What is LongWoF-Bench?

Evaluates verifiable long-workflow tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following, with…

Released
2026-08-24
Evaluates
Agents, Agents & Tool Use, Code, Code generation, Mathematics & Formal Science, Reasoning, Software & AI Compute
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

LongWoF-Bench paper

No reported scores are on record for this benchmark yet.

Which sources cite LongWoF-Bench?

Related Agents, Agents & Tool Use, Code, Code generation, Mathematics & Formal Science, Reasoning, Software & AI Compute benchmarks