Benchmark

SimpleQA

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains…

Modality
text
Categories
reasoning, factuality, general
Openness
unknown
Source
llm_stats
Reported scores
47

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
DeepSeek-R1-0528DeepSeek0.9232025-05-28evidence
DeepSeek-V3DeepSeek0.2492024-12-25evidence
DeepSeek-V3.1DeepSeek0.9342025-01-10evidence
DeepSeek-V3.2-ExpDeepSeek0.9712025-09-29evidence
DeepSeek-V4-Flash-0423DeepSeek0.2892026-04-23evidence
DeepSeek-V4-Flash-MaxDeepSeek0.3412026-04-23evidence
DeepSeek-V4-Pro-MaxDeepSeek0.5792026-04-23evidence
ERNIE 4.5Baidu0.0182025-06-25evidence
ERNIE 5.0Baidu0.752026-01-22evidence
Gemini 2.0 Flash-LiteGoogle0.2172025-02-05evidence
Gemini 2.5 FlashGoogle0.2692025-05-20evidence
Gemini 2.5 Flash-LiteGoogle0.1072025-06-17evidence
Gemini 2.5 ProGoogle0.5082025-05-20evidence
Gemini 2.5 Pro Preview 06-05Google0.542025-06-05evidence
Gemini 3 FlashGoogle0.6872025-12-17evidence
Gemini 3 ProGoogle0.7212025-11-18evidence
Gemini 3.1 Flash-LiteGoogle0.4332026-03-03evidence
Gemma 3 12BGoogle0.0632025-03-12evidence
Gemma 3 1BGoogle0.0222025-03-12evidence
Gemma 3 27BGoogle0.12025-03-12evidence
Gemma 3 4BGoogle0.042025-03-12evidence
GPT-4.5OpenAI0.6252025-02-27evidence
GPT-4oOpenAI0.3822024-08-06evidence
Grok 4 FastxAI0.952025-08-28evidence
Kimi K2 BaseMoonshot AI0.3532025-07-11evidence
Kimi K2 InstructMoonshot AI0.312025-07-11evidence
Kimi K2-Instruct-0905Moonshot AI0.312025-09-05evidence
MiniMax M1 40KMiniMax0.1792025-06-16evidence
MiniMax M1 80KMiniMax0.1852025-06-16evidence
Mistral Large 3 (675B Base)Mistral0.2382025-12-04evidence
Mistral Large 3 (675B Instruct 2512 Eagle)Mistral0.2382025-12-04evidence
Mistral Large 3 (675B Instruct 2512 NVFP4)Mistral0.2382025-12-04evidence
Mistral Large 3 (675B Instruct 2512)Mistral0.2382025-12-04evidence
Mistral Small 3.1 24B InstructMistral0.10432025-03-17evidence
Mistral Small 3.2 24B InstructMistral0.1212025-06-20evidence
o1OpenAI0.472024-12-17evidence
o1-previewOpenAI0.4242024-09-12evidence
o3-miniOpenAI0.152025-01-30evidence
Phi 4Microsoft0.032024-12-12evidence
Qwen3 VL 235B A22B InstructQwen0.5192025-09-22evidence
Qwen3 VL 235B A22B ThinkingQwen0.4442025-09-22evidence
Qwen3 VL 30B A3B InstructQwen0.272025-09-22evidence
Qwen3 VL 30B A3B ThinkingQwen0.2392025-09-22evidence
Qwen3 VL 32B ThinkingQwen0.5542025-09-22evidence
Qwen3 VL 4B InstructQwen0.482025-09-22evidence
Qwen3 VL 8B ThinkingQwen0.4962025-09-22evidence
Qwen3-235B-A22B-Instruct-2507Qwen0.5432025-07-22evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.