HealthBench Professional · independent analysis

Read the system, the slice and the score.

Independent HealthBench Professional analysis: task selection, clinician workflows, system harness comparisons and an interactive length-adjustment calculator.

Independent analysis by Arcophos · updated

One model, three system conditions

Study details ↗

April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task. Compare the whole system: the base model, a browsing harness and ChatGPT for Clinicians.

Length-adjusted rubric score × 100 · points
GPT-5.4 condition
050100
Score
ChatGPT for CliniciansProduct harness59
baseBase-model condition48.1
browsingSimple browsing harness45.8

Paper-reported system comparison. The original figure reports 95% intervals; exact endpoints are not transcribed here. Harness differences are part of the tested system. Sections 5.1 and 5.3; Figure 7. [2]

15,079candidate examples
525selected difficult tasks
[2]
236

Care consult [2]

142

Writing and documentation [2]

147

Medical research [2]

Counts sum to 525. Difficulty slices differ within each use case, so use-case performance does not isolate intrinsic task difficulty. This site explains the April 2026 study; the dedicated leaderboard carries current published results.

An original analytical tool

Inspect the answer-length adjustment

Inspect the arithmetic

Enter an illustrative raw rubric score and final-answer character count. The calculator applies the April 2026 paper’s length adjustment before any benchmark-level clipping.

Illustrative response / not a reported result

Coefficient: 0.0000294 score units per character. Reference length: 2,000 characters. Inputs are a hypothetical response, not a model’s published average.

Per-response adjusted contribution50.00%

+0.00 percentage points from length adjustment

raw − 0.0000294 × (characters − 2000)

The raw score is converted from percent to a fraction before applying the coefficient. This per-response contribution is not clipped; the published overall procedure clips the mean after aggregation. The arithmetic does not establish that a shorter response is clinically better.

This is a score calculation, not an evaluation run or an answer-length recommendation. The coefficient was fitted over a limited length regime; aggregate scores require all examples. [2][3]

The benchmark in detail

All dossiers →
Dossier01

April 2026 paper v1; length-adjusted primary score

HealthBench Professional ↗

Professional task scores depend on sampling, response length and the harness.

UnitClinician task / conversationMeasureLength-adjusted mean rubric score

What we examine

HealthBench Professional combines physician-authored tasks, deliberate difficulty selection and a length-adjusted rubric score. Those choices make its number useful for a particular kind of comparison. This publication explains the sampling, shows historical same-model system results and makes the adjustment formula interactive. Read the benchmark dossier, inspect the calculator and use the guides to separate measured performance from broader clinical claims. Arcophos publishes independent analysis; OpenAI and the original research team created the benchmark.

Selection
Understand how difficulty enrichment changes the meaning of the average.
System
Distinguish a base model from browsing and product-harness conditions.
Scoring
Trace signed rubric points through final-answer length adjustment.

Analysis & interpretation

All analyses →

Questions, answered

What does HealthBench Professional evaluate?

Selected clinician-facing conversation tasks across care consult, writing/documentation and medical research, graded with physician-written rubrics and a length-adjusted primary score.

How does the length adjustment work?

In fractional units, subtract 0.0000294 times final-answer characters minus 2,000 from the raw rubric score. Aggregate clipping happens after averaging; the calculator displays illustrative pre-aggregation values.

Does 59 points mean 59% of clinical tasks are correct?

No. It is an aggregate length-adjusted rubric measurement on deliberately selected tasks. It is neither binary task accuracy nor an estimated routine clinical success rate.

Can the paper evaluation be reproduced exactly from public code?

The official card says the paper uses an internal implementation and no official external equivalent is released. Public reference settings support related evaluation, but do not establish exact reproduction of the internal harness.

Is this an official OpenAI publication?

No. This site is independent Arcophos analysis. The original paper, release and dataset card identify the benchmark’s creators and remain the primary sources.

Prepare a comparison worksheet

Working tool / saved on this device

Document a professional-system comparison

Interactive worksheet

Record the benchmark and harness choices behind a reported result. This worksheet does not execute the benchmark or certify a clinical system.

Specify the selected tasks

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗