Benchmark

HiddenMath

Google DeepMind's internal mathematical reasoning benchmark that introduces novel problems not encountered during model training to evaluate true…

Modality
text
Categories
math, reasoning
Openness
unknown
Source
llm_stats
Reported scores
13

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Gemini 1.5 FlashGoogle0.4722024-05-01evidence
Gemini 1.5 Flash 8BGoogle0.3282024-03-15evidence
Gemini 1.5 ProGoogle0.522024-05-01evidence
Gemini 2.0 FlashGoogle0.632024-12-01evidence
Gemini 2.0 Flash-LiteGoogle0.5532025-02-05evidence
Gemma 3 12BGoogle0.5452025-03-12evidence
Gemma 3 1BGoogle0.1582025-03-12evidence
Gemma 3 27BGoogle0.6032025-03-12evidence
Gemma 3 4BGoogle0.432025-03-12evidence
Gemma 3n E2B InstructedGoogle0.2772025-06-26evidence
Gemma 3n E2B Instructed LiteRT (Preview)Google0.2772025-05-20evidence
Gemma 3n E4B InstructedGoogle0.3772025-06-26evidence
Gemma 3n E4B Instructed LiteRT PreviewGoogle0.3772025-05-20evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.