Benchmark

HealthBench Professional

HealthBench Professional evaluates model capability and safety for clinician use cases using real clinician-style chats and physician-authored grading…

Modality
text
Categories
healthcare
Openness
unknown
Source
llm_stats
Reported scores
9

Reported scores

llm_stats

ModelOrganizationReported valueReportedEvidence
Claude Fable 5Anthropic0.662026-06-09evidence
Claude Opus 4.8Anthropic0.5582026-05-28evidence
Claude Opus 5Anthropic0.5982026-07-24evidence
Claude Sonnet 5Anthropic0.5782026-06-30evidence
GPT-5.5 InstantOpenAI0.3842026-05-05evidence
GPT-5.6 LunaOpenAI0.5572026-07-09evidence
GPT-5.6 SolOpenAI0.6052026-07-09evidence
GPT-5.6 TerraOpenAI0.5772026-07-09evidence
MAI-Thinking-1Microsoft0.352026-06-02evidence

Scores are partitioned by the source that reported them and are never merged into a single cross-source ranking, because the sources measure different things and say so.