Benchmark
VQAv2
VQAv2 is a balanced Visual Question Answering dataset that addresses language bias by providing complementary images for each question, forcing models to…
- Modality
- multimodal
- Categories
- multimodal, reasoning, image_to_text, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 3
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Llama 3.2 90B Instruct | Meta | 0.781 | 2024-09-25 | evidence |
| Pixtral Large | Mistral | 0.809 | 2024-11-18 | evidence |
| Pixtral-12B | Mistral | 0.786 | 2024-09-17 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.