Benchmark
CSimpleQA
Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions. It contains…
- Modality
- text
- Categories
- language, general
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 8
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| DeepSeek-V3 | DeepSeek | 0.648 | 2024-12-25 | evidence |
| DeepSeek-V4-Flash-0423 | DeepSeek | 0.732 | 2026-04-23 | evidence |
| DeepSeek-V4-Flash-Max | DeepSeek | 0.789 | 2026-04-23 | evidence |
| DeepSeek-V4-Pro-Max | DeepSeek | 0.844 | 2026-04-23 | evidence |
| Kimi K2 Base | Moonshot AI | 0.776 | 2025-07-11 | evidence |
| Kimi K2 Instruct | Moonshot AI | 0.784 | 2025-07-11 | evidence |
| Qwen3 VL 235B A22B Instruct | Qwen | 0.834 | 2025-09-22 | evidence |
| Qwen3-235B-A22B-Instruct-2507 | Qwen | 0.843 | 2025-07-22 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.