Benchmark

Multi-Challenge

MultiChallenge is a realistic multi-turn conversation evaluation benchmark that challenges frontier LLMs across four key categories: instruction retention…

Modality
text
Categories
reasoning, communication
Openness
unknown
Source
llm_stats
Reported scores
29

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT-4.1OpenAI0.3832025-04-14evidence
GPT-4.1 miniOpenAI0.3582025-04-14evidence
GPT-4.1 nanoOpenAI0.152025-04-14evidence
GPT-4.5OpenAI0.4382025-02-27evidence
GPT-4oOpenAI0.4032024-08-06evidence
GPT-5OpenAI0.6962025-08-07evidence
Kimi K2 InstructMoonshot AI0.5412025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.5412025-09-05evidence
MAI-Thinking-1Microsoft0.532026-06-02evidence
MiniMax M1 40KMiniMax0.4472025-06-16evidence
MiniMax M1 80KMiniMax0.4472025-06-16evidence
Nemotron 3 Nano (30B A3B)NVIDIA0.3852025-12-15evidence
Nemotron 3 Super (120B A12B)NVIDIA0.55232026-03-11evidence
Nemotron 3 Ultra (550B A55B)NVIDIA0.6382026-06-04evidence
Nova 2 LiteAmazon0.7662025-12-02evidence
Nova 2 OmniAmazon0.7552025-12-02evidence
Nova 2 ProAmazon0.7772025-12-02evidence
o3OpenAI0.6042025-04-16evidence
o3-miniOpenAI0.3992025-01-30evidence
o4-miniOpenAI0.432025-04-16evidence
Qwen3.5-0.8BQwen0.1892026-03-02evidence
Qwen3.5-122B-A10BQwen0.6152026-02-24evidence
Qwen3.5-27BQwen0.6082026-02-24evidence
Qwen3.5-2BQwen0.3372026-03-02evidence
Qwen3.5-35B-A3BQwen0.62026-02-24evidence
Qwen3.5-397B-A17BQwen0.6762026-02-16evidence
Qwen3.5-4BQwen0.492026-03-02evidence
Qwen3.5-9BQwen0.5452026-03-02evidence
Step3-VL-10BStepFun0.6262026-01-15evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.