Benchmark

POPE

Polling-based Object Probing Evaluation (POPE) is a benchmark for evaluating object hallucination in Large Vision-Language Models (LVLMs). POPE addresses…

Modality
multimodal
Categories
multimodal, safety, vision
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
LFM2.5-VL-3BLiquid AI0.8872026-08-12evidence
Phi-3.5-vision-instructMicrosoft0.8612024-08-23evidence
Phi-4-multimodal-instructMicrosoft0.8562025-02-01evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.