Benchmark

C-Eval

C-Eval is a comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese…

Modality
text
Categories
reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
18

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V3DeepSeek0.8652024-12-25evidence
ERNIE 4.5Baidu0.4072025-06-25evidence
Kimi K2 BaseMoonshot AI0.9252025-07-11evidence
Kimi-k1.5Moonshot AI0.8832025-01-20evidence
MiMo-V2.5-ProXiaomi0.9152026-04-27evidence
Qwen2 72B InstructQwen0.8382024-07-23evidence
Qwen2 7B InstructQwen0.7722024-07-23evidence
Qwen3.5-0.8BQwen0.5052026-03-02evidence
Qwen3.5-122B-A10BQwen0.9192026-02-24evidence
Qwen3.5-27BQwen0.9052026-02-24evidence
Qwen3.5-2BQwen0.7322026-03-02evidence
Qwen3.5-35B-A3BQwen0.9022026-02-24evidence
Qwen3.5-397B-A17BQwen0.932026-02-16evidence
Qwen3.5-4BQwen0.8512026-03-02evidence
Qwen3.5-9BQwen0.8822026-03-02evidence
Qwen3.6 PlusQwen0.9332026-04-02evidence
Qwen3.6-27BQwen0.9142026-04-21evidence
Qwen3.6-35B-A3BQwen0.92026-04-16evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.