Benchmark

AlignBench

AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories…

Modality
text
Categories
math, reasoning, roleplay, language, general, creativity, writing
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-V2.5DeepSeek0.8042024-05-08evidence
Qwen2 7B InstructQwen0.7212024-07-23evidence
Qwen2.5 72B InstructQwen0.8162024-09-19evidence
Qwen2.5 7B InstructQwen0.7332024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.