Who this is for
This resource is written for teams evaluating clinician-facing AI systems and interpreting professional-workflow benchmark results. Its scope is benchmark-specific analysis: what the published tasks measure, how scores are produced, and what the evidence can support.
- Selection. Understand how difficulty enrichment changes the meaning of the average.
- System. Distinguish a base model from browsing and product-harness conditions.
- Scoring. Trace signed rubric points through final-answer length adjustment.
Ownership and independence
HealthBench Professional Guide is published by Arcophos. Arcophos also publishes other healthcare AI evaluation resources linked from this site. Those links are disclosed as related resources, not independent endorsements.
The site is not the official home of any third-party benchmark, regulator, hospital, or model developer discussed in its guides. Benchmark names and publication titles identify their respective authors’ work. Our checklists and interpretations are editorial material.
How to read this site
Start with a benchmark dossier, follow its original references, and use the analytical explorer to inspect the tasks or scoring assumptions. The planning worksheet remains available as a secondary tool. Each guide states a direct answer, develops its reasoning, and links supporting references.
These materials do not establish clinical efficacy, patient benefit, or regulatory compliance for any system. Our editorial method explains the boundaries.
Your worksheet data
Worksheet selections are stored in this browser on this device. They are not sent to Arcophos. You can reset selections at any time or download them as a text file. This site has no analytics or advertising scripts.
Contact and corrections
For a correction, include the page URL, the specific claim, and a link to the original evidence. Contact Arcophos ↗