Benchmark

DROP

DROP (Discrete Reasoning Over Paragraphs) is a reading comprehension benchmark requiring discrete reasoning over paragraph content. It contains…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
30

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude 3 HaikuAnthropic0.7842024-03-13evidence
Claude 3 OpusAnthropic0.8312024-02-29evidence
Claude 3 SonnetAnthropic0.7892024-02-29evidence
Claude 3.5 HaikuAnthropic0.8312024-10-22evidence
Claude 3.5 SonnetAnthropic0.8712024-06-21evidence
Claude 3.5 SonnetAnthropic0.8712024-10-22evidence
DeepSeek-V3DeepSeek0.9162024-12-25evidence
ERNIE 4.5Baidu0.2862025-06-25evidence
Gemini 1.5 ProGoogle0.7492024-05-01evidence
Gemma 3n E2BGoogle0.5392025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.5392025-05-20evidence
Gemma 3n E4BGoogle0.6082025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.6082025-05-20evidence
GPT-3.5 TurboOpenAI0.7022023-03-21evidence
GPT-4OpenAI0.8092023-06-13evidence
GPT-4 TurboOpenAI0.862024-04-09evidence
GPT-4oOpenAI0.8342024-05-13evidence
GPT-4o miniOpenAI0.7972024-07-18evidence
Granite 3.3 8B BaseIBM0.36142025-04-16evidence
Granite 3.3 8B InstructIBM0.59362025-04-16evidence
IBM Granite 4.0 Tiny PreviewIBM0.4622025-05-02evidence
Llama 3.1 405B InstructMeta0.8482024-07-23evidence
Llama 3.1 70B InstructMeta0.7962024-07-23evidence
Llama 3.1 8B InstructMeta0.5952024-07-23evidence
LongCat-Flash-ChatMeituan0.79062025-08-29evidence
MiMo-V2.5-ProXiaomi0.8632026-04-27evidence
Nova LiteAmazon0.8022024-11-20evidence
Nova MicroAmazon0.7932024-11-20evidence
Nova ProAmazon0.8542024-11-20evidence
Phi 4Microsoft0.7552024-12-12evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.