Benchmark

Codeforces

An Elo estimate against human contestants, not a fixed dataset. The problem set moves continuously and rating conversion differs by vendor.

Released
2010-02-19
Categories
coding
Openness
unknown
Source
model_reports
Reported scores
3

Reported scores

model_reports

ModelOrganizationReported valueReportedEvidence
DeepSeek-V4-FlashDeepSeek3052.02026-06-25evidence
DeepSeek-V4-ProDeepSeek3206.02026-04-22evidence
Gemma 4 (31B)Google2150.02026-03-11evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.