Benchmark
OCRBench_V2
OCRBench v2: Enhanced large-scale bilingual benchmark for evaluating Large Multimodal Models on visual text localization and reasoning with 10,000…
- Modality
- multimodal
- Categories
- image_to_text, vision
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 7
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Nova 2 Lite | Amazon | 0.561 | 2025-12-02 | evidence |
| Nova 2 Omni | Amazon | 0.582 | 2025-12-02 | evidence |
| Nova 2 Pro | Amazon | 0.645 | 2025-12-02 | evidence |
| Qwen2.5-Omni-7B | Qwen | 0.578 | 2025-03-27 | evidence |
| Qwen3.7-Plus | Qwen | 0.671 | 2026-05-31 | evidence |
| Seed 2.1 Pro | ByteDance | 0.632 | 2026-06-24 | evidence |
| Seed 2.1 Turbo | ByteDance | 0.628 | 2026-06-24 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.