Benchmark

CSimpleQA

Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions. It contains…

Modality
text
Categories
language, general
Openness
unknown
Source
llm_stats
Reported scores
8

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3DeepSeek0.6482024-12-25evidence
DeepSeek-V4-Flash-0423DeepSeek0.7322026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.7892026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.8442026-04-23evidence
Kimi K2 BaseMoonshot AI0.7762025-07-11evidence
Kimi K2 InstructMoonshot AI0.7842025-07-11evidence
Qwen3 VL 235B A22B InstructQwen0.8342025-09-22evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.8432025-07-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.