Benchmark

OmniDocBench 1.5

OmniDocBench 1.5 is a comprehensive benchmark for evaluating multimodal large language models on document understanding tasks, including OCR, document…

Modality
multimodal
Categories
multimodal, reasoning, structured_output, vision
Openness
unknown
Source
llm_stats
Reported scores
18

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DiffusionGemma 26B-A4BGoogle0.3192026-06-10evidence
Gemini 3 FlashGoogle0.1212025-12-17evidence
Gemini 3 ProGoogle0.1152025-11-18evidence
Gemma 4 12BGoogle0.1642026-05-23evidence
GPT-5.4OpenAI0.8912026-03-05evidence
GPT-5.4 miniOpenAI0.87372026-03-17evidence
GPT-5.4 nanoOpenAI0.75812026-03-17evidence
GPT-5.5 InstantOpenAI0.8752026-05-05evidence
Kimi K2.5Moonshot AI0.8882026-01-27evidence
MiniMax M3MiniMax0.9162026-06-01evidence
Muse Glimmer-30BMeta0.7582026-08-10evidence
Qwen3.5-122B-A10BQwen0.8982026-02-24evidence
Qwen3.5-27BQwen0.8892026-02-24evidence
Qwen3.5-35B-A3BQwen0.8932026-02-24evidence
Qwen3.6 PlusQwen0.9122026-04-02evidence
Qwen3.6-35B-A3BQwen0.8992026-04-16evidence
Qwen3.7-PlusQwen0.9142026-05-31evidence
Qwen3.8-27BQwen0.9112026-08-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.