{"publication":"HealthBench Professional Guide","url":"https://healthbenchprofessional.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"healthbench-professional","name":"HealthBench Professional","shortName":"HealthBench Professional","version":"April 2026 paper v1; length-adjusted primary score","creators":"Rebecca Soskin Hicks and colleagues · OpenAI","paperDate":"2026-04-30","headline":"Professional task scores depend on sampling, response length and the harness.","summary":"HealthBench Professional focuses on clinician-facing conversations and deliberately selects difficult examples. Its primary metric further adjusts rubric scores for answer length. Those design choices give the benchmark a useful stress-testing role while limiting how its headline number should be generalized to routine work. We analyze the use-case and difficulty mix, explain the adjustment algebra, and separate the model from its product harness. Results below are from the original April 2026 study. They are not new evaluations, current rankings, or estimates of patient benefit.","task":{"input":"A single- or multi-turn physician-authored conversation ending with a clinician request.","output":"The next assistant response, graded with physician-written rubric items.","unit":"Clinician task / conversation","setting":"525 selected tasks; default GPT-5.4 low-reasoning grader; primary score adjusted for final-answer length."},"dataOrigin":"Physicians testing ChatGPT for Clinicians in good-faith and adversarial modes; reviewed and sampled for difficulty, rather than a random patient cohort.","facts":[{"label":"Selected tasks","value":"525","detail":"Final evaluation set.","sourceIds":["hbp-paper"],"locator":"Sections 2 and 3.4"},{"label":"Candidate examples","value":"15,079","detail":"Pool before review and stratified selection.","sourceIds":["hbp-paper"],"locator":"Section 3.4"},{"label":"Physician contributors","value":"190","detail":"Contributors to creation and review.","sourceIds":["hbp-paper"],"locator":"Section 3.1"},{"label":"Use cases","value":"3","detail":"Care consult, writing/documentation, medical research.","sourceIds":["hbp-paper"],"locator":"Section 2"},{"label":"Main-comparison samples","value":"8 per task","detail":"Study averages repeated samples within examples.","sourceIds":["hbp-paper"],"locator":"Section 4.2"},{"label":"Adjustment center","value":"2,000 characters","detail":"Final-answer characters; thinking tokens are excluded.","sourceIds":["hbp-paper"],"locator":"Section 4.1"}],"metric":{"name":"Length-adjusted mean rubric score","description":"Normalize signed rubric points, subtract the length adjustment for each answer, then average and clip the mean to [0,1]. Display times 100.","formula":"sᵢ,adjusted = sᵢ − 0.0000294 × (charactersᵢ − 2000); score = 100 × clip(mean(sᵢ,adjusted), 0, 1)","direction":"Higher is better under the published length-adjusted rubric.","comparability":"Fix grader, answer-length definition, reasoning effort, sample count, task selection and system harness. The published evaluator is internal.","sourceIds":["hbp-paper","hbp-card"]},"workflow":[{"label":"Supply the clinician conversation","detail":"Evaluate the next response for the selected task.","sourceIds":["hbp-paper"]},{"label":"Run the specified system","detail":"Main comparisons use highest available reasoning effort, default verbosity and eight samples per example.","sourceIds":["hbp-paper"]},{"label":"Grade rubric items","detail":"GPT-5.4 at low reasoning decides criterion satisfaction in the published protocol.","sourceIds":["hbp-paper"]},{"label":"Adjust for response length","detail":"Use the published slope and center on final-answer characters.","sourceIds":["hbp-paper"]},{"label":"Average and clip","detail":"Aggregate adjusted scores with the study’s repeated-sample treatment.","sourceIds":["hbp-paper"]}],"slices":[{"label":"Care consult","value":236,"unit":"tasks","detail":"Use-case total, Appendix Table 2","sourceIds":["hbp-paper"]},{"label":"Writing and documentation","value":142,"unit":"tasks","detail":"Use-case total, Appendix Table 2","sourceIds":["hbp-paper"]},{"label":"Medical research","value":147,"unit":"tasks","detail":"Use-case total, Appendix Table 2","sourceIds":["hbp-paper"]}],"sliceTitle":"The three use cases","sliceNote":"Counts sum to 525. Difficulty slices differ within each use case, so use-case performance does not isolate intrinsic task difficulty.","results":[{"id":"hbp-harness","title":"Same model, different system conditions","metric":"Length-adjusted rubric score × 100","unit":"points","lower":0,"upper":100,"scope":"April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task.","sourceIds":["hbp-paper"],"locator":"Sections 5.1 and 5.3; Figure 7","rows":[{"label":"GPT-5.4 · ChatGPT for Clinicians","value":59,"display":"59","detail":"Product harness"},{"label":"GPT-5.4 · base","value":48.1,"display":"48.1","detail":"Base-model condition"},{"label":"GPT-5.4 · browsing","value":45.8,"display":"45.8","detail":"Simple browsing harness"}],"note":"Paper-reported system comparison. The original figure reports 95% intervals; exact endpoints are not transcribed here. Harness differences are part of the tested system."}],"analysis":[{"heading":"The final sample emphasizes failures.","evidence":"The authors selected 525 examples from 15,079 candidates with difficulty enrichment.","interpretation":"An aggregate score is a stress-test measurement, not an estimate of the frequency of errors in ordinary clinical use.","sourceIds":["hbp-paper"]},{"heading":"Use case and difficulty mix are entangled.","evidence":"Writing has 91 red-teaming difficult tasks of 142; research has 7 of 147.","interpretation":"Comparing their raw subscores alone cannot show which activity is inherently harder. The selected case mix also changed.","sourceIds":["hbp-paper"]},{"heading":"Longer answers must earn their adjustment.","evidence":"The slope is 0.0000294 raw-score units per final-answer character.","interpretation":"An added 1,000 characters requires 2.94 more raw score points to break even. This is a calculation, not advice to shorten clinical content.","sourceIds":["hbp-paper"]},{"heading":"A harness comparison is not a model swap.","evidence":"The reported overall scores differ across three GPT-5.4 system conditions.","interpretation":"Attribute the comparison to the full configuration. The model name alone cannot reproduce the result.","sourceIds":["hbp-paper","hbp-card"]}],"limitations":[{"title":"Internal evaluation harness","detail":"The official card states no official external implementation is released; the public reference is not the complete internal evaluator.","sourceIds":["hbp-card"]},{"title":"Selected hard cases","detail":"Difficulty ratings were assigned against recent OpenAI models and used in sampling, affecting the evaluation distribution.","sourceIds":["hbp-paper"]},{"title":"Length adjustment has a fitting range","detail":"The coefficient was estimated in a limited response-length regime; very long answers may behave differently.","sourceIds":["hbp-paper"]},{"title":"Clinical scope remains bounded","detail":"The paper does not cover every workflow, including institution-specific and EHR-integrated work.","sourceIds":["hbp-paper"]}],"access":{"status":"Official benchmark data are available; paper evaluation uses an internal implementation.","license":"Official dataset-card metadata: MIT.","restrictions":"Authors request that examples not be posted in text or images online. This analysis publishes only aggregate metadata and formulas.","url":"https://huggingface.co/datasets/openai/healthbench-professional","sourceIds":["hbp-card"]},"sourceIds":["hbp-paper","hbp-card","hbp-release","hb-code"]}],"explorer":{"kind":"length-adjustment","title":"Inspect the answer-length adjustment","intro":"Enter an illustrative raw rubric score and final-answer character count. The calculator applies the April 2026 paper’s length adjustment before any benchmark-level clipping.","caution":"This is a score calculation, not an evaluation run or an answer-length recommendation. The coefficient was fitted over a limited length regime; aggregate scores require all examples.","sourceIds":["hbp-paper","hbp-card"],"rows":[],"parameters":[{"key":"slope","value":0.0000294,"label":"Raw-score adjustment per character","sourceIds":["hbp-paper"]},{"key":"center","value":2000,"label":"Reference final-answer length in characters","sourceIds":["hbp-paper"]}]},"references":[{"id":"hb-code","title":"HealthBench reference evaluator","organization":"OpenAI","url":"https://github.com/openai/simple-evals/blob/652c89d0ca9df547706735883097e9537d40dc47/healthbench_eval.py","note":"Pinned reference scoring implementation, including positive-point normalization.","locator":"calculate_score; calculate_length_adjusted_score; aggregate results","version":"Commit 652c89d0ca9df547706735883097e9537d40dc47"},{"id":"hbp-paper","title":"HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats","organization":"Rebecca Soskin Hicks and colleagues · OpenAI","url":"https://arxiv.org/html/2604.27470v1","note":"Primary sampling, length adjustment, model/harness comparison and composition tables.","locator":"Sections 3–6 and 8; Appendix Tables 1–3","version":"arXiv v1 · 2026-04-30"},{"id":"hbp-card","title":"HealthBench Professional official dataset card","organization":"OpenAI","url":"https://huggingface.co/datasets/openai/healthbench-professional","note":"MIT metadata, field definitions, contamination request and internal-evaluator limitation.","locator":"Dataset card; implementation note","version":"Accessed 2026-09-28"},{"id":"hbp-release","title":"Making ChatGPT better for clinicians","organization":"OpenAI","url":"https://openai.com/index/making-chatgpt-better-for-clinicians/","note":"Dated announcement introducing the benchmark; separate from the subsequent arXiv version.","locator":"Continuing to evaluate and strengthen model health performance and safety","version":"Published 2026-04-22"}]}