The final sample emphasizes failures.
The authors selected 525 examples from 15,079 candidates with difficulty enrichment. [2]
An aggregate score is a stress-test measurement, not an estimate of the frequency of errors in ordinary clinical use.
Independent benchmark analysis / April 2026 paper v1; length-adjusted primary score
HealthBench Professional focuses on clinician-facing conversations and deliberately selects difficult examples. Its primary metric further adjusts rubric scores for answer length. Those design choices give the benchmark a useful stress-testing role while limiting how its headline number should be generalized to routine work. We analyze the use-case and difficulty mix, explain the adjustment algebra, and separate the model from its product harness. Results below are from the original April 2026 study. They are not new evaluations, current rankings, or estimates of patient benefit.
01 / What is being tested?
Data origin. Physicians testing ChatGPT for Clinicians in good-faith and adversarial modes; reviewed and sampled for difficulty, rather than a random patient cohort. [2][3][4][1]
Final evaluation set.
Sections 2 and 3.4 [2]Pool before review and stratified selection.
Section 3.4 [2]Contributors to creation and review.
Section 3.1 [2]Care consult, writing/documentation, medical research.
Section 2 [2]Study averages repeated samples within examples.
Section 4.2 [2]Final-answer characters; thinking tokens are excluded.
Section 4.1 [2]Evaluate the next response for the selected task.
[2]Main comparisons use highest available reasoning effort, default verbosity and eight samples per example.
[2]GPT-5.4 at low reasoning decides criterion satisfaction in the published protocol.
[2]Use the published slope and center on final-answer characters.
[2]Aggregate adjusted scores with the study’s repeated-sample treatment.
[2]Dataset anatomy
Use-case total, Appendix Table 2
Use-case total, Appendix Table 2
Use-case total, Appendix Table 2
Counts sum to 525. Difficulty slices differ within each use case, so use-case performance does not isolate intrinsic task difficulty. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better under the published length-adjusted rubric.
Normalize signed rubric points, subtract the length adjustment for each answer, then average and clip the mean to [0,1]. Display times 100.
sᵢ,adjusted = sᵢ − 0.0000294 × (charactersᵢ − 2000); score = 100 × clip(mean(sᵢ,adjusted), 0, 1)
Fix grader, answer-length definition, reasoning effort, sample count, task selection and system harness. The published evaluator is internal. [2][3]
03 / Measured evidence
Paper-reported results / selected rows
April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task.
Paper-reported system comparison. The original figure reports 95% intervals; exact endpoints are not transcribed here. Harness differences are part of the tested system.
Source: Sections 5.1 and 5.3; Figure 7 [2]
04 / Our original analysis
The authors selected 525 examples from 15,079 candidates with difficulty enrichment. [2]
An aggregate score is a stress-test measurement, not an estimate of the frequency of errors in ordinary clinical use.
Writing has 91 red-teaming difficult tasks of 142; research has 7 of 147. [2]
Comparing their raw subscores alone cannot show which activity is inherently harder. The selected case mix also changed.
The slope is 0.0000294 raw-score units per final-answer character. [2]
An added 1,000 characters requires 2.94 more raw score points to break even. This is a calculation, not advice to shorten clinical content.
05 / Scope of the evidence
The official card states no official external implementation is released; the public reference is not the complete internal evaluator. [3]
Difficulty ratings were assigned against recent OpenAI models and used in sampling, affecting the evaluation distribution. [2]
The coefficient was estimated in a limited response-length regime; very long answers may behave differently. [2]
The paper does not cover every workflow, including institution-specific and EHR-integrated work. [2]
Evidence trail
OpenAI. Pinned reference scoring implementation, including positive-point normalization.
Rebecca Soskin Hicks and colleagues · OpenAI. Primary sampling, length adjustment, model/harness comparison and composition tables.
OpenAI. MIT metadata, field definitions, contamination request and internal-evaluator limitation.
OpenAI. Dated announcement introducing the benchmark; separate from the subsequent arXiv version.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.