Benchmark

SWE-Lancer

A benchmark for evaluating large language models on real-world freelance software engineering tasks from Upwork. Contains over 1,400 tasks valued at $1…

Modality
text
Categories
reasoning, code
Openness
unknown
Source
llm_stats
Reported scores
4

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-4.5OpenAI0.3732025-02-27evidence
GPT-4oOpenAI0.3262024-08-06evidence
GPT-5.1 CodexOpenAI0.6632025-11-19evidence
o3-miniOpenAI0.182025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.