Evidence / original sources

A claim should lead somewhere.

Follow the references behind HealthBench Professional Guide. Each entry explains its role, links to the original publication, and identifies the guides that use it.

Content updated September 28, 2026. Inclusion does not imply endorsement.

02

Rebecca Soskin Hicks and colleagues · OpenAI

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats ↗

Primary sampling, length adjustment, model/harness comparison and composition tables.

Version: arXiv v1 · 2026-04-30

Evidence locator: Sections 3–6 and 8; Appendix Tables 1–3

arxiv.org