Benchmark

IFBench

Instruction Following Benchmark evaluating model's ability to follow complex instructions

Modality
text
Categories
instruction_following, general
Openness
unknown
Source
llm_stats
Reported scores
34

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Command A+Cohere0.742026-05-20evidence
GPT OSS 120B HighOpenAI0.6952025-08-05evidence
Hermes 3 70BNous Research0.81212024-08-15evidence
Inkling-SmallThinking Machines Lab0.8222026-07-30evidence
K-EXAONE-236B-A23BLG AI Research0.6732025-12-31evidence
LFM2.5-2.6BLiquid AI0.59172026-08-04evidence
LFM2.5-VL-3BLiquid AI0.2582026-08-12evidence
MAI-Code-1-FlashMicrosoft0.752026-06-02evidence
MAI-Thinking-1Microsoft0.692026-06-02evidence
Mercury 2Inception0.712026-02-24evidence
MiniMax M2.1MiniMax0.72025-12-23evidence
Mistral Medium 3.5Mistral0.692026-04-29evidence
Mistral Small 4Mistral0.482026-03-16evidence
Muse Glimmer-30BMeta0.772026-08-10evidence
Nemotron 3 Super (120B A12B)NVIDIA0.72562026-03-11evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.8172026-06-04evidence
Nemotron 3.5 Lightning (30B A3B)NVIDIA0.71882026-08-11evidence
Nova 2 LiteAmazon0.7082025-12-02evidence
Nova 2 OmniAmazon0.6872025-12-02evidence
Nova 2 ProAmazon0.8022025-12-02evidence
Nova 2 SonicAmazon0.3752025-12-02evidence
Qwen3.5-0.8BQwen0.212026-03-02evidence
Qwen3.5-122B-A10BQwen0.7612026-02-24evidence
Qwen3.5-27BQwen0.7652026-02-24evidence
Qwen3.5-2BQwen0.4132026-03-02evidence
Qwen3.5-35B-A3BQwen0.7022026-02-24evidence
Qwen3.5-397B-A17BQwen0.7652026-02-16evidence
Qwen3.5-4BQwen0.5922026-03-02evidence
Qwen3.5-9BQwen0.6452026-03-02evidence
Qwen3.6 PlusQwen0.7422026-04-02evidence
Qwen3.7 MaxQwen0.7912026-05-19evidence
Qwen3.7-PlusQwen0.7912026-05-31evidence
Qwen3.8 MaxQwen0.8282026-08-02evidence
Qwen3.8-27BQwen0.7952026-08-14evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.