Benchmark

FACTS Grounding

A benchmark evaluating language models' ability to generate factually accurate and well-grounded responses based on long-form input context, comprising…

Modality
text
Categories
reasoning, factuality, grounding
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 2.0 FlashGoogle0.8362024-12-01evidence
Gemini 2.0 Flash-LiteGoogle0.8362025-02-05evidence
Gemini 2.5 FlashGoogle0.8532025-05-20evidence
Gemini 2.5 Flash-LiteGoogle0.8412025-06-17evidence
Gemini 2.5 Pro Preview 06-05Google0.8782025-06-05evidence
Gemini 3 FlashGoogle0.6192025-12-17evidence
Gemini 3 ProGoogle0.7052025-11-18evidence
Gemini 3.1 Flash-LiteGoogle0.4062026-03-03evidence
Gemma 3 12BGoogle0.7582025-03-12evidence
Gemma 3 1BGoogle0.3642025-03-12evidence
Gemma 3 27BGoogle0.7492025-03-12evidence
Gemma 3 4BGoogle0.7012025-03-12evidence
GLM-5V-TurboZ.ai0.5862026-04-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.