April 2026 paper v1; length-adjusted primary score
HealthBench Professional ↗
Professional task scores depend on sampling, response length and the harness.
HealthBench Professional · independent analysis
Independent HealthBench Professional analysis: task selection, clinician workflows, system harness comparisons and an interactive length-adjustment calculator.
April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task. Compare the whole system: the base model, a browsing harness and ChatGPT for Clinicians.
Paper-reported system comparison. The original figure reports 95% intervals; exact endpoints are not transcribed here. Harness differences are part of the tested system. Sections 5.1 and 5.3; Figure 7. [2]
Counts sum to 525. Difficulty slices differ within each use case, so use-case performance does not isolate intrinsic task difficulty. This site explains the April 2026 study; the dedicated leaderboard carries current published results.
An original analytical tool
Enter an illustrative raw rubric score and final-answer character count. The calculator applies the April 2026 paper’s length adjustment before any benchmark-level clipping.
Illustrative response / not a reported result
Coefficient: 0.0000294 score units per character. Reference length: 2,000 characters. Inputs are a hypothetical response, not a model’s published average.
+0.00 percentage points from length adjustment
raw − 0.0000294 × (characters − 2000)
The raw score is converted from percent to a fraction before applying the coefficient. This per-response contribution is not clipped; the published overall procedure clips the mean after aggregation. The arithmetic does not establish that a shorter response is clinically better.
This is a score calculation, not an evaluation run or an answer-length recommendation. The coefficient was fitted over a limited length regime; aggregate scores require all examples. [2][3]
April 2026 paper v1; length-adjusted primary score
Professional task scores depend on sampling, response length and the harness.
HealthBench Professional combines physician-authored tasks, deliberate difficulty selection and a length-adjusted rubric score. Those choices make its number useful for a particular kind of comparison. This publication explains the sampling, shows historical same-model system results and makes the adjustment formula interactive. Read the benchmark dossier, inspect the calculator and use the guides to separate measured performance from broader clinical claims. Arcophos publishes independent analysis; OpenAI and the original research team created the benchmark.
Apply the published coefficient in the correct units and keep example-level values separate from aggregate clipping.
Read difficulty enrichment and use-case composition before generalizing a HealthBench Professional score.
Interpret the original same-model system experiment without attributing every change to the base model.
Selected clinician-facing conversation tasks across care consult, writing/documentation and medical research, graded with physician-written rubrics and a length-adjusted primary score.
In fractional units, subtract 0.0000294 times final-answer characters minus 2,000 from the raw rubric score. Aggregate clipping happens after averaging; the calculator displays illustrative pre-aggregation values.
No. It is an aggregate length-adjusted rubric measurement on deliberately selected tasks. It is neither binary task accuracy nor an estimated routine clinical success rate.
The official card says the paper uses an internal implementation and no official external equivalent is released. Public reference settings support related evaluation, but do not establish exact reproduction of the internal harness.
No. This site is independent Arcophos analysis. The original paper, release and dataset card identify the benchmark’s creators and remain the primary sources.
Working tool / saved on this device
Record the benchmark and harness choices behind a reported result. This worksheet does not execute the benchmark or certify a clinical system.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗