Benchmark

Aider-Polyglot

A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models…

Modality
text
Categories
general, code
Openness
unknown
Source
llm_stats
Reported scores
22

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1-0528DeepSeek0.7162025-05-28evidence
DeepSeek-V3DeepSeek0.4962024-12-25evidence
DeepSeek-V3.1DeepSeek0.6842025-01-10evidence
DeepSeek-V3.2-ExpDeepSeek0.7452025-09-29evidence
Gemini 2.5 FlashGoogle0.6192025-05-20evidence
Gemini 2.5 Flash-LiteGoogle0.2672025-06-17evidence
Gemini 2.5 ProGoogle0.7652025-05-20evidence
Gemini 2.5 Pro Preview 06-05Google0.8222025-06-05evidence
GPT-4.1OpenAI0.5162025-04-14evidence
GPT-4.1 miniOpenAI0.3472025-04-14evidence
GPT-4.1 nanoOpenAI0.0982025-04-14evidence
GPT-4oOpenAI0.3072024-08-06evidence
GPT-5OpenAI0.882025-08-07evidence
Kimi K2 InstructMoonshot AI0.62025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.62025-09-05evidence
Magistral MediumMistral0.4712025-06-10evidence
o3OpenAI0.8132025-04-16evidence
o3-miniOpenAI0.6672025-01-30evidence
o4-miniOpenAI0.6892025-04-16evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.5732025-07-22evidence
Qwen3-Coder 480B A35B InstructQwen0.6182025-01-31evidence
Qwen3-Next-80B-A3B-InstructQwen0.4982025-09-10evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.