Benchmark

PhiBench

PhiBench is an internal benchmark designed to evaluate diverse skills and reasoning abilities of language models, covering a wide range of tasks including…

Modality
text
Categories
math, reasoning, general
Openness
unknown
Source
llm_stats
Reported scores
3

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Phi 4Microsoft0.5622024-12-12evidence
Phi 4 ReasoningMicrosoft0.7062025-04-30evidence
Phi 4 Reasoning PlusMicrosoft0.7422025-04-30evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.