Independent benchmark analysis / April 2026 paper v1; length-adjusted primary score

HealthBench Professional

Professional task scores depend on sampling, response length and the harness.

HealthBench Professional focuses on clinician-facing conversations and deliberately selects difficult examples. Its primary metric further adjusts rubric scores for answer length. Those design choices give the benchmark a useful stress-testing role while limiting how its headline number should be generalized to routine work. We analyze the use-case and difficulty mix, explain the adjustment algebra, and separate the model from its product harness. Results below are from the original April 2026 study. They are not new evaluations, current rankings, or estimates of patient benefit.

01 / What is being tested?

The task, before the score.

input
A single- or multi-turn physician-authored conversation ending with a clinician request.
output
The next assistant response, graded with physician-written rubric items.
unit
Clinician task / conversation
setting
525 selected tasks; default GPT-5.4 low-reasoning grader; primary score adjusted for final-answer length.

Data origin. Physicians testing ChatGPT for Clinicians in good-faith and adversarial modes; reviewed and sampled for difficulty, rather than a random patient cohort. [2][3][4][1]

Selected tasks
525

Final evaluation set.

Sections 2 and 3.4 [2]
Candidate examples
15,079

Pool before review and stratified selection.

Section 3.4 [2]
Physician contributors
190

Contributors to creation and review.

Section 3.1 [2]
Use cases
3

Care consult, writing/documentation, medical research.

Section 2 [2]
Main-comparison samples
8 per task

Study averages repeated samples within examples.

Section 4.2 [2]
Adjustment center
2,000 characters

Final-answer characters; thinking tokens are excluded.

Section 4.1 [2]
  1. 01

    Supply the clinician conversation

    Evaluate the next response for the selected task.

    [2]
  2. 02

    Run the specified system

    Main comparisons use highest available reasoning effort, default verbosity and eight samples per example.

    [2]
  3. 03

    Grade rubric items

    GPT-5.4 at low reasoning decides criterion satisfaction in the published protocol.

    [2]
  4. 04

    Adjust for response length

    Use the published slope and center on final-answer characters.

    [2]
  5. 05

    Average and clip

    Aggregate adjusted scores with the study’s repeated-sample treatment.

    [2]

Dataset anatomy

The three use cases

Care consult

Use-case total, Appendix Table 2

236 tasks[2]
Writing and documentation

Use-case total, Appendix Table 2

142 tasks[2]
Medical research

Use-case total, Appendix Table 2

147 tasks[2]

Counts sum to 525. Difficulty slices differ within each use case, so use-case performance does not isolate intrinsic task difficulty. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Length-adjusted mean rubric score

Higher is better under the published length-adjusted rubric.

Normalize signed rubric points, subtract the length adjustment for each answer, then average and clip the mean to [0,1]. Display times 100.

Scoring definition

sᵢ,adjusted = sᵢ − 0.0000294 × (charactersᵢ − 2000); score = 100 × clip(mean(sᵢ,adjusted), 0, 1)

Fix grader, answer-length definition, reasoning effort, sample count, task selection and system harness. The published evaluator is internal. [2][3]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Same model, different system conditions

April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task.

Length-adjusted rubric score × 100 · points
050100
Reported
GPT-5.4 · ChatGPT for CliniciansProduct harness
59
GPT-5.4 · baseBase-model condition
48.1
GPT-5.4 · browsingSimple browsing harness
45.8

Paper-reported system comparison. The original figure reports 95% intervals; exact endpoints are not transcribed here. Harness differences are part of the tested system.

Source: Sections 5.1 and 5.3; Figure 7 [2]

04 / Our original analysis

What follows from the design?

01

The final sample emphasizes failures.

Published evidence

The authors selected 525 examples from 15,079 candidates with difficulty enrichment. [2]

Our interpretation

An aggregate score is a stress-test measurement, not an estimate of the frequency of errors in ordinary clinical use.

02

Use case and difficulty mix are entangled.

Published evidence

Writing has 91 red-teaming difficult tasks of 142; research has 7 of 147. [2]

Our interpretation

Comparing their raw subscores alone cannot show which activity is inherently harder. The selected case mix also changed.

03

Longer answers must earn their adjustment.

Published evidence

The slope is 0.0000294 raw-score units per final-answer character. [2]

Our interpretation

An added 1,000 characters requires 2.94 more raw score points to break even. This is a calculation, not advice to shorten clinical content.

04

A harness comparison is not a model swap.

Published evidence

The reported overall scores differ across three GPT-5.4 system conditions. [2][3]

Our interpretation

Attribute the comparison to the full configuration. The model name alone cannot reproduce the result.

05 / Scope of the evidence

Where this benchmark stops.

Internal evaluation harness

The official card states no official external implementation is released; the public reference is not the complete internal evaluator. [3]

Selected hard cases

Difficulty ratings were assigned against recent OpenAI models and used in sampling, affecting the evaluation distribution. [2]

Length adjustment has a fitting range

The coefficient was estimated in a limited response-length regime; very long answers may behave differently. [2]

Clinical scope remains bounded

The paper does not cover every workflow, including institution-specific and EHR-integrated work. [2]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Official benchmark data are available; paper evaluation uses an internal implementation.
License
Official dataset-card metadata: MIT.
Conditions
Authors request that examples not be posted in text or images online. This analysis publishes only aggregate metadata and formulas.
[3]

Evidence trail

Read the originals.

  1. HealthBench reference evaluator ↗

    OpenAI. Pinned reference scoring implementation, including positive-point normalization.

  2. HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats ↗

    Rebecca Soskin Hicks and colleagues · OpenAI. Primary sampling, length adjustment, model/harness comparison and composition tables.

  3. HealthBench Professional official dataset card ↗

    OpenAI. MIT metadata, field definitions, contamination request and internal-evaluator limitation.

  4. Making ChatGPT better for clinicians ↗

    OpenAI. Dated announcement introducing the benchmark; separate from the subsequent arXiv version.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗