Benchmark

Internal API instruction following (hard)

Internal API instruction following (hard) benchmark - specific documentation not found in official sources

Modality
text
Categories
structured_output, general
Openness
unknown
Source
llm_stats
Reported scores
7

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-4.1OpenAI0.4912025-04-14evidence
GPT-4.1 miniOpenAI0.4512025-04-14evidence
GPT-4.1 nanoOpenAI0.3162025-04-14evidence
GPT-4.5OpenAI0.542025-02-27evidence
GPT-4oOpenAI0.2922024-08-06evidence
GPT-5OpenAI0.642025-08-07evidence
o3-miniOpenAI0.52025-01-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.