Benchmark

HealthBench Hard

A challenging variation of HealthBench that evaluates large language models' performance and safety in healthcare through 5,000 multi-turn conversations…

Modality
text
Categories
healthcare
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT OSS 120BOpenAI0.32025-08-05evidence
GPT OSS 20BOpenAI0.1082025-08-05evidence
GPT-5OpenAI0.0162025-08-07evidence
GPT-5.3 ChatOpenAI0.2592026-03-04evidence
GPT-5.5 InstantOpenAI0.2292026-05-05evidence
GPT-5.6 LunaOpenAI0.322026-07-09evidence
GPT-5.6 SolOpenAI0.3312026-07-09evidence
GPT-5.6 TerraOpenAI0.3272026-07-09evidence
Muse SparkMeta0.4282026-04-08evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.