Benchmark
TruthfulQA
TruthfulQA is a benchmark to measure whether language models are truthful in generating answers to questions. It comprises 817 questions that span 38…
- Modality
- text
- Categories
- legal, reasoning, finance, general, healthcare
- Openness
- unknown
- Source
- llm_stats
- Reported scores
- 18
Reported scores
llm_stats
| Model | Organization | Reported value | Reported | Evidence |
|---|---|---|---|---|
| Command R+ | Cohere | 0.563 | 2024-08-30 | evidence |
| Granite 3.3 8B Base | IBM | 0.5215 | 2025-04-16 | evidence |
| Granite 3.3 8B Instruct | IBM | 0.6686 | 2025-04-16 | evidence |
| Hermes 3 70B | Nous Research | 0.6329 | 2024-08-15 | evidence |
| IBM Granite 4.0 Tiny Preview | IBM | 0.581 | 2025-05-02 | evidence |
| Jamba 1.5 Large | AI21 Labs | 0.583 | 2024-08-22 | evidence |
| Jamba 1.5 Mini | AI21 Labs | 0.541 | 2024-08-22 | evidence |
| Llama 3.1 Nemotron 70B Instruct | NVIDIA | 0.5863 | 2024-10-01 | evidence |
| MAI-Thinking-1 | Microsoft | 0.88 | 2026-06-02 | evidence |
| Mistral NeMo Instruct | Mistral | 0.503 | 2024-07-18 | evidence |
| Phi 4 Mini | Microsoft | 0.664 | 2025-02-01 | evidence |
| Phi-3.5-mini-instruct | Microsoft | 0.64 | 2024-08-23 | evidence |
| Phi-3.5-MoE-instruct | Microsoft | 0.775 | 2024-08-23 | evidence |
| Qwen2 72B Instruct | Qwen | 0.548 | 2024-07-23 | evidence |
| Qwen2.5 14B Instruct | Qwen | 0.584 | 2024-09-19 | evidence |
| Qwen2.5 32B Instruct | Qwen | 0.578 | 2024-09-19 | evidence |
| Qwen2.5-Coder 32B Instruct | Qwen | 0.542 | 2024-09-19 | evidence |
| Qwen2.5-Coder 7B Instruct | Qwen | 0.506 | 2024-09-19 | evidence |
Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.