Benchmark

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from…

Released
2024-03-12
Modality
text
Categories
reasoning, general, code
Openness
unknown
Source
llm_stats
Reported scores
75

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek R1 Distill Llama 70BDeepSeek0.5752025-01-20evidence
DeepSeek R1 Distill Llama 8BDeepSeek0.3962025-01-20evidence
DeepSeek R1 Distill Qwen 1.5BDeepSeek0.1692025-01-20evidence
DeepSeek R1 Distill Qwen 14BDeepSeek0.5312025-01-20evidence
DeepSeek R1 Distill Qwen 32BDeepSeek0.5722025-01-20evidence
DeepSeek R1 Distill Qwen 7BDeepSeek0.3762025-01-20evidence
DeepSeek R1 ZeroDeepSeek0.52025-01-20evidence
DeepSeek-R1-0528DeepSeek0.7332025-05-28evidence
DeepSeek-V3DeepSeek0.3762024-12-25evidence
DeepSeek-V3 0324DeepSeek0.4922025-03-25evidence
DeepSeek-V3.1DeepSeek0.5642025-01-10evidence
DeepSeek-V3.2DeepSeek0.8332025-12-01evidence
DeepSeek-V3.2 (Thinking)DeepSeek0.8332025-12-01evidence
DeepSeek-V3.2-ExpDeepSeek0.7412025-09-29evidence
DeepSeek-V4-Flash-0423DeepSeek0.8842026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.9162026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.9352026-04-23evidence
Gemini 2.0 FlashGoogle0.3512024-12-01evidence
Gemini 2.5 Flash-LiteGoogle0.3372025-06-17evidence
Gemini 2.5 Pro Preview 06-05Google0.692025-06-05evidence
Gemini DiffusionGoogle0.3092025-05-20evidence
Gemma 3 12BGoogle0.2462025-03-12evidence
Gemma 3 1BGoogle0.0192025-03-12evidence
Gemma 3 27BGoogle0.2972025-03-12evidence
Gemma 3 4BGoogle0.1262025-03-12evidence
Gemma 3n E2B InstructedGoogle0.1322025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.1322025-05-20evidence
Gemma 3n E4B InstructedGoogle0.1322025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.1322025-05-20evidence
GLM-4.5Z.ai0.7292025-07-28evidence
GLM-4.5-AirZ.ai0.7072025-07-28evidence
Grok 4 FastxAI0.82025-08-28evidence
Grok-3xAI0.7942025-02-17evidence
Grok-3 MinixAI0.8042025-02-17evidence
Grok-4xAI0.792025-07-09evidence
Grok-4 HeavyxAI0.7942025-07-09evidence
Kimi K2-Instruct-0905Moonshot AI0.5372025-09-05evidence
Llama 3.1 Nemotron Ultra 253B v1NVIDIA0.66312025-04-07evidence
Llama 4 MaverickMeta0.4342025-04-05evidence
Llama 4 ScoutMeta0.3282025-04-05evidence
LongCat-Flash-ChatMeituan0.48022025-08-29evidence
LongCat-Flash-ThinkingMeituan0.7942025-09-22evidence
LongCat-Flash-Thinking-2601Meituan0.8282026-01-14evidence
Magistral MediumMistral0.5032025-06-10evidence
Magistral Small 2506Mistral0.5132025-06-10evidence
Mercury 2Inception0.672026-02-24evidence
Min istral 3 (3B Reasoning 2512)Mistral0.5482025-12-04evidence
MiniMax M1 40KMiniMax0.6232025-06-16evidence
MiniMax M1 80KMiniMax0.652025-06-16evidence
MiniMax M2MiniMax0.832025-10-27evidence
MiniMax M2.1MiniMax0.782025-12-23evidence
Ministral 3 (14B Reasoning 2512)Mistral0.6462025-12-04evidence
Ministral 3 (8B Reasoning 2512)Mistral0.6162025-12-04evidence
Mistral Large 3 (675B Base)Mistral0.3442025-12-04evidence
Mistral Large 3 (675B Instruct 2512 Eagle)Mistral0.3442025-12-04evidence
Mistral Large 3 (675B Instruct 2512 NVFP4)Mistral0.3442025-12-04evidence
Mistral Large 3 (675B Instruct 2512)Mistral0.3442025-12-04evidence
Mistral Small 4Mistral0.6362026-03-16evidence
Nemotron 3 Super (120B A12B)NVIDIA0.81192026-03-11evidence
Nemotron Nano 9B v2NVIDIA0.7112025-08-18evidence
Nova 2 LiteAmazon0.712025-12-02evidence
Nova 2 ProAmazon0.7462025-12-02evidence
Phi 4 ReasoningMicrosoft0.5382025-04-30evidence
Phi 4 Reasoning PlusMicrosoft0.5312025-04-30evidence
Qwen2 7B InstructQwen0.2662024-07-23evidence
Qwen2.5 72B InstructQwen0.5552024-09-19evidence
Qwen2.5 7B InstructQwen0.2872024-09-19evidence
Qwen2.5-Coder 32B InstructQwen0.3142024-09-19evidence
Qwen2.5-Coder 7B InstructQwen0.1822024-09-19evidence
Qwen3 235B A22BQwen0.7072025-04-29evidence
Qwen3 30B A3BQwen0.6262025-04-29evidence
Qwen3 32BQwen0.6572025-04-29evidence
QwQ-32BQwen0.6342025-03-05evidence
QwQ-32B-PreviewQwen0.52024-11-28evidence
Solar Pro 4Upstage0.8782026-08-06evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.