Benchmark
EvalPlus
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to…
- Modality
- text
- Categories
- reasoning, code
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 4
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Kimi K2 Base | Moonshot AI | 0.803 | 2025-07-11 | evidence |
| Qwen2 72B Instruct | Qwen | 0.79 | 2024-07-23 | evidence |
| Qwen2 7B Instruct | Qwen | 0.703 | 2024-07-23 | evidence |
| Qwen3 235B A22B | Qwen | 0.776 | 2025-04-29 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.