Benchmark

LongDS-Bench

What is LongDS-Bench?

LongDS-Bench evaluates long-horizon, multi-turn data analysis tasks where agents must maintain, update, restore, and compose evolving analytical states. It…

Released
2026-05-28
Evaluates
General AI, Language & Knowledge, cs.LG
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

LongDS-Bench paper and code

No reported scores are on record for this benchmark yet.

Which sources cite LongDS-Bench?

Related General AI, Language & Knowledge, cs.LG benchmarks