The short answer
A system can change while its base model remains the same. HealthBench Professional makes that distinction visible by comparing GPT-5.4 under base, browsing and ChatGPT for Clinicians conditions. This guide treats the experiment as a system comparison, preserves its historical protocol and explains what additional artifacts would be needed to reproduce or extend the result.
Define the system beyond the model name
A harness determines how a model is prompted, which tools it can use and how a workflow is carried out around it. Those choices can alter the information and opportunities available before the final response. Naming only the base model therefore leaves part of the evaluated object unspecified.
For a comparison receipt, list the base model and the surrounding conditions separately. If a product harness is not fully described or publicly released, mark that as a reproducibility limit. The result can still be reported as a published measurement, but it should not be presented as an experiment another team can exactly reconstruct from a model identifier alone.
Read the published overall comparison
The April 2026 paper reports length-adjusted overall scores of 59.0 for GPT-5.4 in ChatGPT for Clinicians, 48.1 for base GPT-5.4 and 45.8 for GPT-5.4 with browsing. Our panel reproduces that historical comparison and its named conditions. It is not a current ranking and contains no new Arcophos evaluation runs.
The measured difference belongs to the full system condition. It does not show that a different base model was used, and it does not isolate one particular prompt or tool as the cause. A causal explanation for an individual component would need a comparison designed to separate that component from the other changes.
Inspect the task-specific pattern
The paper also examines use cases and source slices. That matters because an overall gain can be concentrated in particular types of work. A system that handles adversarial documentation requests differently may change the mean without improving every kind of clinician interaction to the same degree.
Before transferring a headline result to an application, identify which tasks the application actually performs. Then ask whether the reported advantage appears in the relevant portion of the benchmark and whether that portion has adequate evidence. A broad product comparison should not erase the heterogeneity that the study itself analyzes.
Keep human responses in their protocol
The study’s physician baseline consists of responses written for the benchmark tasks under stated conditions, including specialty matching, internet access and no time limit. That is a specific response-writing comparison. It does not directly measure bedside care, team coordination or how quickly clinicians work with an assistant in practice.
If a report includes the human baseline, attach those conditions to it. The intended inference should concern the scored response task. A phrase such as outperforming physicians becomes ambiguous when separated from what the physicians were asked to produce and what the rubric actually measured.
Separate a public reference from the internal evaluator
The official data card states that the paper uses an internal evaluation implementation and does not release an official external equivalent. The public reference can lower the barrier to related evaluation, but matching selected grader settings does not establish that every aspect of a product harness or internal protocol has been reproduced.
For a new study, publish the configuration you can actually specify and label departures from the source protocol. Preserve raw and adjusted scores, response lengths and task identifiers. The most useful conclusion names the tested system, measured task distribution and comparison conditions, allowing readers to understand the result without guessing which parts came from the model and which came from its surrounding workflow.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats ↗Rebecca Soskin Hicks and colleagues · OpenAI. Primary sampling, length adjustment, model/harness comparison and composition tables.
- HealthBench Professional official dataset card ↗OpenAI. MIT metadata, field definitions, contamination request and internal-evaluator limitation.
- Making ChatGPT better for clinicians ↗OpenAI. Dated announcement introducing the benchmark; separate from the subsequent arXiv version.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.