Benchmark
HumanEval+
Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code…
- Modality
- text
- Categories
- reasoning
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 10
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| ERNIE 4.5 | Baidu | 0.25 | 2025-06-25 | evidence |
| Granite 3.3 8B Base | IBM | 0.8609 | 2025-04-16 | evidence |
| Granite 3.3 8B Instruct | IBM | 0.8609 | 2025-04-16 | evidence |
| IBM Granite 4.0 Tiny Preview | IBM | 0.783 | 2025-05-02 | evidence |
| MiMo-V2.5-Pro | Xiaomi | 0.756 | 2026-04-27 | evidence |
| Phi 4 | Microsoft | 0.828 | 2024-12-12 | evidence |
| Phi 4 Reasoning | Microsoft | 0.929 | 2025-04-30 | evidence |
| Phi 4 Reasoning Plus | Microsoft | 0.923 | 2025-04-30 | evidence |
| Qwen2.5 14B Instruct | Qwen | 0.512 | 2024-09-19 | evidence |
| Qwen2.5 32B Instruct | Qwen | 0.524 | 2024-09-19 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.