Benchmark

EvalPlus

A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to…

Modality
text
Categories
reasoning, code
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Kimi K2 BaseMoonshot AI0.8032025-07-11evidence
Qwen2 72B InstructQwen0.792024-07-23evidence
Qwen2 7B InstructQwen0.7032024-07-23evidence
Qwen3 235B A22BQwen0.7762025-04-29evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.