Benchmark

HealthBench

An open-source benchmark for measuring performance and safety of large language models in healthcare, consisting of 5,000 multi-turn conversations…

Modality
text
Categories
healthcare
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
GPT OSS 120BOpenAI0.5762025-08-05evidence
GPT OSS 20BOpenAI0.4252025-08-05evidence
GPT-5.3 ChatOpenAI0.5412026-03-04evidence
GPT-5.5 InstantOpenAI0.5142026-05-05evidence
GPT-5.6 LunaOpenAI0.5582026-07-09evidence
GPT-5.6 SolOpenAI0.572026-07-09evidence
GPT-5.6 TerraOpenAI0.572026-07-09evidence
Kimi K2-Thinking-0905Moonshot AI0.582025-09-05evidence
Qwen3.8 MaxQwen0.6022026-08-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.