Benchmark

CodeForces

A competitive programming benchmark using problems from the CodeForces platform. The benchmark evaluates code generation capabilities of LLMs on…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1-0528DeepSeek0.64332025-05-28evidence
DeepSeek-V3.1DeepSeek0.6972025-01-10evidence
DeepSeek-V3.2DeepSeek0.7952025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.7952025-12-01evidence
DeepSeek-V3.2-ExpDeepSeek0.7072025-09-29evidence
DeepSeek-V3.2-SpecialeDeepSeek0.92025-12-01evidence
DeepSeek-V4-Flash-0423DeepSeek0.93872026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek1.02026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek1.02026-04-23evidence
DiffusionGemma 26B-A4BGoogle0.47632026-06-10evidence
Gemma 4 12BGoogle0.5532026-05-23evidence
GPT OSS 120BOpenAI0.8212025-08-05evidence
GPT OSS 20BOpenAI0.74332025-08-05evidence
Qwen3 32BQwen0.6592025-04-29evidence
Qwen3.5-122B-A10BQwen0.8512026-02-24evidence
Qwen3.5-27BQwen0.8072026-02-24evidence
Qwen3.5-35B-A3BQwen0.8222026-02-24evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.