How HealthBench Professional adjusts for answer length
Apply the published coefficient in the correct units and keep example-level values separate from aggregate clipping.
Original analysis / HealthBench Professional Guide
HealthBench Professional combines physician-authored tasks, deliberate difficulty selection and a length-adjusted rubric score. Those choices make its number useful for a particular kind of comparison. This publication explains the sampling, shows historical same-model system results and makes the adjustment formula interactive. Read the benchmark dossier, inspect the calculator and use the guides to separate measured performance from broader clinical claims. Arcophos publishes independent analysis; OpenAI and the original research team created the benchmark.
Apply the published coefficient in the correct units and keep example-level values separate from aggregate clipping.
Read difficulty enrichment and use-case composition before generalizing a HealthBench Professional score.
Interpret the original same-model system experiment without attributing every change to the base model.