Benchmark

TruthfulQA

TruthfulQA is a benchmark to measure whether language models are truthful in generating answers to questions. It comprises 817 questions that span 38…

Modality
text
Categories
legal, reasoning, finance, general, healthcare
Openness
unknown
Source
llm_stats
Reported scores
18

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Command R+Cohere0.5632024-08-30evidence
Granite 3.3 8B BaseIBM0.52152025-04-16evidence
Granite 3.3 8B InstructIBM0.66862025-04-16evidence
Hermes 3 70BNous Research0.63292024-08-15evidence
IBM Granite 4.0 Tiny PreviewIBM0.5812025-05-02evidence
Jamba 1.5 LargeAI21 Labs0.5832024-08-22evidence
Jamba 1.5 MiniAI21 Labs0.5412024-08-22evidence
Llama 3.1 Nemotron 70B InstructNVIDIA0.58632024-10-01evidence
MAI-Thinking-1Microsoft0.882026-06-02evidence
Mistral NeMo InstructMistral0.5032024-07-18evidence
Phi 4 MiniMicrosoft0.6642025-02-01evidence
Phi-3.5-mini-instructMicrosoft0.642024-08-23evidence
Phi-3.5-MoE-instructMicrosoft0.7752024-08-23evidence
Qwen2 72B InstructQwen0.5482024-07-23evidence
Qwen2.5 14B InstructQwen0.5842024-09-19evidence
Qwen2.5 32B InstructQwen0.5782024-09-19evidence
Qwen2.5-Coder 32B InstructQwen0.5422024-09-19evidence
Qwen2.5-Coder 7B InstructQwen0.5062024-09-19evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.