Benchmark

MMVetGPT4Turbo

MM-Vet evaluation using GPT-4 Turbo for scoring. This variant of MM-Vet examines large multimodal models on complicated multimodal tasks requiring…

Modality
multimodal
Categories
math, multimodal, reasoning, spatial_reasoning, general, vision
Openness
unknown
Source
llm_stats
Reported scores
1

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Qwen2-VL-72B-InstructQwen0.742024-08-29evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.