Benchmark

LongEarth-Bench

What is LongEarth-Bench?

LongEarth-Bench evaluates vision-language models on long-horizon Earth observation reasoning, with ~120k QA samples from 117k images, sequences avg 15.14…

Released
2026-08-13
Evaluates
General AI, Multimodal, Reasoning
Openness
unknown
Importer
Claire Radar
Review status
unreviewed
Reported scores
0

Source provenance

LongEarth-Bench paper

No reported scores are on record for this benchmark yet.

Which sources cite LongEarth-Bench?

Related General AI, Multimodal, Reasoning benchmarks