Benchmark

HumanEval+

Enhanced version of HumanEval that extends the original test cases by 80x using EvalPlus framework for rigorous evaluation of LLM-synthesized code…

Modality
text
Categories
reasoning
Openness
unknown
Source
llm_stats
Reported scores
10

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
ERNIE 4.5Baidu0.252025-06-25evidence
Granite 3.3 8B BaseIBM0.86092025-04-16evidence
Granite 3.3 8B InstructIBM0.86092025-04-16evidence
IBM Granite 4.0 Tiny PreviewIBM0.7832025-05-02evidence
MiMo-V2.5-ProXiaomi0.7562026-04-27evidence
Phi 4Microsoft0.8282024-12-12evidence
Phi 4 ReasoningMicrosoft0.9292025-04-30evidence
Phi 4 Reasoning PlusMicrosoft0.9232025-04-30evidence
Qwen2.5 14B InstructQwen0.5122024-09-19evidence
Qwen2.5 32B InstructQwen0.5242024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.