Benchmark

MTVQA

MTVQA (Multilingual Text-Centric Visual Question Answering) is the first benchmark featuring high-quality human expert annotations across 9 diverse…

Modality
multimodal
Categories
multimodal, text-to-image, vision
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen2-VL-72B-InstructQwen0.3092024-08-29evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.