Benchmark
HumanEval
Saturated. Retained because open-weight cards still report it.
- Released
- 2021-07-07
- Categories
- coding
- Openness
- unknown
- Source
- model_reports
- Reported scores
- 2
Reported scores
model_reports
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Gemini 1.5 Pro | 84.1 | 2024-03-08 | evidence | |
| Qwen2.5-Coder-32B-Instruct | Qwen | 92.7 | 2024-09-18 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.