The short answer
HealthBench Professional reports a primary score adjusted for final-answer length. The adjustment is a mathematical component of the evaluation, not a recommendation for how long a clinical response should be. Our calculator exposes that component with illustrative inputs. This guide explains its units, its effect on comparisons and the boundary between one adjusted answer score and a complete benchmark result.
Begin with the raw rubric fraction
The underlying rubric score sums the signed points of met criteria and divides by the available positive points for the example. This produces a fraction before the length adjustment. It can be negative when penalties exceed earned credit. Displaying that fraction on a 100-point scale is a separate presentation choice.
Keep those units explicit when implementing the formula. A coefficient defined for a fraction cannot be subtracted unchanged from a percentage-scale number. Our calculator accepts a readable score input and applies the published arithmetic in fractional units before converting back to score points for display.
Apply the published slope and center
The paper subtracts 0.0000294 multiplied by the difference between final-answer characters and 2,000. Answers at the center receive no change. Longer answers receive a negative adjustment; shorter answers receive a positive adjustment. The length concerns the final answer, not hidden reasoning tokens.
For a hypothetical raw fraction of 0.60 and an answer of 3,000 characters, the adjustment is 0.0294. The resulting example score is 0.5706, or 57.06 points. These numbers are an arithmetic illustration, not a reported model result or a released benchmark example.
Interpret break-even changes carefully
The formula implies that adding 1,000 characters requires 2.94 additional raw score points to leave the adjusted score unchanged. That is a derived property of the measurement. It can help explain why two configurations with different raw scores receive similar adjusted scores, even before inspecting their response content.
It does not establish that removing text improves an answer. A shorter response may lose required information, while a longer response may add necessary detail. The adjustment only describes how the benchmark treats length once the grader has assigned the underlying rubric score. Content quality and character count remain separate inputs.
Preserve the fitting-range limitation
The authors estimate the coefficient using controlled verbosity comparisons and a limited answer-length regime. Applying a linear rule far outside the region used to estimate it may have a different interpretation. The paper explicitly discusses tradeoffs for very long responses, so the coefficient should not be described as a universal law of clinical usefulness.
When reporting an evaluation, retain answer-length distributions alongside raw and adjusted scores. This allows a reader to see whether a difference is driven by rubric performance, response length or both. It also makes an unusually long-output configuration visible rather than hiding it inside the final average.
Clip the benchmark mean, not every example
After adjustment, the published method averages example scores and clips the aggregate mean to the interval from zero to one. A single example’s pre-aggregation value should therefore remain visible even if it is negative or exceeds one. Premature clipping would implement a different calculation.
The calculator stops at the example-level arithmetic and explains that boundary. A complete result still needs the full task set, grader, generation protocol and aggregation procedure. The official data card also distinguishes the internal evaluation implementation from the public reference. Matching the formula is necessary for comparability, but it is not the same as reproducing the whole published system.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats ↗Rebecca Soskin Hicks and colleagues · OpenAI. Primary sampling, length adjustment, model/harness comparison and composition tables.
- HealthBench Professional official dataset card ↗OpenAI. MIT metadata, field definitions, contamination request and internal-evaluator limitation.
- HealthBench reference evaluator ↗OpenAI. Pinned reference scoring implementation, including positive-point normalization.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.