Benchmark

FrontierMath

A benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians, covering most major…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
17

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-5OpenAI0.2632025-08-07evidence
GPT-5 miniOpenAI0.2212025-08-07evidence
GPT-5 nanoOpenAI0.0962025-08-07evidence
GPT-5.1OpenAI0.2672025-11-13evidence
GPT-5.1 InstantOpenAI0.2672025-11-12evidence
GPT-5.1 ThinkingOpenAI0.2672025-11-12evidence
GPT-5.2OpenAI0.4032025-12-11evidence
GPT-5.4OpenAI0.4762026-03-05evidence
GPT-5.5OpenAI0.3542026-04-23evidence
GPT-5.5 ProOpenAI0.3962026-04-23evidence
GPT-5.6 LunaOpenAI0.7862026-07-09evidence
GPT-5.6 SolOpenAI0.892026-07-09evidence
GPT-5.6 TerraOpenAI0.8492026-07-09evidence
MAI-Code-1-FlashMicrosoft0.0632026-06-02evidence
o1OpenAI0.0552024-12-17evidence
o3OpenAI0.1582025-04-16evidence
o3-miniOpenAI0.0922025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.