Benchmark

FLEX Benchmark

What is FLEX?

FLEX evaluates expert routing across a 34-task long-horizon multimodal continual instruction-tuning sequence with reduced task-identifying textual fingerprints.

Released
2026-08-02
Evaluates
General AI, Multimodal
Openness
unknown
Importer
Claire Radar
Review status
ai-reviewed
Reported scores
0

Source provenance

FLEX paper and code

No reported scores are on record for this benchmark yet.

Which sources cite FLEX?

Related General AI, Multimodal benchmarks