Benchmark

MM IF-Eval

A challenging multimodal instruction-following benchmark that includes both compose-level constraints for output responses and perception-level…

Modality
multimodal
Categories
multimodal, reasoning, structured_output
Openness
unknown
Source
llm_stats
Reported scores
2

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
LFM2.5-VL-3BLiquid AI0.6062026-08-12evidence
Pixtral-12BMistral0.5272024-09-17evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.