When a company builds the exam, supplies the graders, and tops the results, the headline writes itself — but the verification does not. OpenAI's June 2026 "health intelligence" update makes exactly that kind of claim, and the first thing worth examining is what actually shipped.
Health Intelligence: The Specific Improvements Delivered
The core change is a distribution shift, not a new product. On June 18, 2026, OpenAI rolled health-tuned behavior into GPT-5.5 Instant — the model released in May 2026 as the new default for free ChatGPT users, subject to usage limits. It replaces GPT-5.3 Instant, the prior default shipped in March 2026. The notable part for developers tracking the stack: OpenAI says this fast, low-cost model matches its pricier "Thinking" models on its own health benchmarks, so stronger health behavior now ships in the free default tier rather than behind a paid gate, according to OpenAI.
OpenAI describes four concrete behavioral changes in the updated model:
- Urgent-care recognition — better at flagging when emergency or urgent care may be needed, with fewer missed red flags.
- Context seeking — asks clarifying questions to gather context before answering.
- Calibrated uncertainty — expresses uncertainty without overstating confidence.
- Localized guidance — tailors medical information to a user's local healthcare context and makes it understandable.
For evidence, OpenAI points to production monitoring rather than a single launch metric. Using privacy-preserving monitors over live health traffic — at the scale of billions of messages weekly — it reports the rate of responses flagged for at least one possible factuality issue fell 71% over two months, per OpenAI. Treat that figure with care: it counts internally flagged issues, and the flagging methodology is not disclosed, so it is not an externally validated error rate.
The scale argument explains why OpenAI is investing here. The company says more than 230 million people ask ChatGPT health and wellness questions each week, with top tasks including interpreting lab results, preparing for appointments, and navigating insurance, according to coverage of the announcement. At that volume, moving the improved behavior into the free default model is the consequential decision — and the basis for the doctor-comparison claims examined in the sections that follow.
HealthBench Professional: How the Physician-Preference Exam Was Constructed

HealthBench Professional is the evaluation that most directly supports the doctor-comparison claim, and it targets clinicians rather than consumers. Submitted to arXiv on April 30, 2026 , the benchmark scores models on 525 clinician-facing tasks selected from 15,079 candidate examples . Those tasks fall into three categories — care consult, writing/documentation, and medical research — so the exam measures the kind of work clinicians actually delegate, not generic Q&A.
The physician network behind it is broad. The benchmark was built with 190 physician contributors across 50 countries, 26 specialties, and 52 languages . Two construction choices push the difficulty up: roughly one-third of the tasks are deliberate physician red-teaming, and difficult examples were enriched approximately 3.5× relative to a natural distribution . The intent is a stress test, not an average-case sample.
The human ceiling is set high on purpose. Physician-written baselines came from specialty-matched doctors given unlimited time and web access, but no AI assistance . That matters for reading any model-beats-doctor headline: the comparison is against a strong, unhurried, reference-equipped human, not a clinician answering under time pressure.
| Design element | HealthBench Professional |
|---|---|
| Tasks (from candidate pool) | 525 of 15,079 |
| Categories | Care consult · Writing/documentation · Medical research |
| Physician contributors | 190 across 50 countries, 26 specialties, 52 languages |
| Red-team share | ~1/3 of tasks |
| Difficulty enrichment | ~3.5× harder examples |
| Human baseline | Specialty-matched MDs, unlimited time + web, no AI |
One scoring detail constrains how the results should be read. The score is not percent accuracy; it uses length-adjusted, rubric-based grading, and the paper explicitly states that absolute scores should not be interpreted as real-world clinical performance rates . In practice, a figure like 59.0 reflects how completely a response satisfies physician-written rubric criteria on a deliberately hard set — a relative preference signal between answers, not a measured clinical success rate. That distinction frames the margin examined next.
The Lab That Built the Rubric Also Won the Exam
Every layer of that margin was produced by a single party. OpenAI authored HealthBench in May 2025 and HealthBench Professional, submitted to arXiv on April 30, 2026 ; it recruited the physician contributors — 262 for the consumer benchmark and 190 for the clinician edition — and it ran its own models against the rubric. Neither paper carries external co-authors. The examiner, the answer key, and the top-scoring candidate share one address.
That alone does not make the results wrong, but it constrains what can be checked. The HealthBench Professional paper states that it uses an internal evaluation implementation and retains a private held-out set . The practical consequence: the headline figures are not fully reproducible from public materials alone. An outside team can read the methodology, but it cannot rerun the exact grading pipeline on the exact data and confirm that GPT-5.4 scored 59.0 against 43.7 for physician-written baselines .
The metric itself is also softer than "beats doctors" implies. The win is a panel preference — physician reviewers rating which response better satisfies rubric criteria, scored on a length-adjusted scale rather than measured against patient outcomes. The hardest red-team slice is a Likert 1-2 band , which is a judgment of answer quality, not a record of whether a patient was correctly treated. As the paper puts it:
"Absolute scores should not be interpreted as real-world clinical performance rates" — HealthBench Professional, OpenAI (source: HealthBench Professional, 2026-04).
What is missing is the part that would settle the question. As of this writing, no independent peer review or external replication of the HealthBench scoring methodology has been published; the benchmark is open , but the clinician-facing scores carry the first-party caveats above. A preference signal generated, graded, and reported by the same lab is a reasonable internal development metric. It is not yet the externally audited evidence that the doctor-versus-AI framing invites readers to assume — a gap the next section weighs against the size of the reported margin.
59.0 vs 43.7: The Physician-Preference Margin

The headline number is a 15.3-point spread: GPT-5.4 running inside ChatGPT for Clinicians scored 59.0 overall on HealthBench Professional versus 43.7 for the physician-written baselines. Base GPT-5.4 — the same model without the clinician-facing tuning — landed at 48.1. That ordering matters: roughly two-thirds of the gap over physicians comes from the specialist RLHF tuning layered on top of the base model, not from the base model alone. The score is length-adjusted rubric scoring, not percent accuracy, and OpenAI states absolute values should not be read as real-world clinical performance rates .
The margin is uneven across the three task categories the benchmark covers. Where the work is generative and structured — drafting documentation — the tuned product nearly doubles the physician baseline. Where it requires synthesis of evidence, the gap narrows to single digits.
| Task category | Tuned product (GPT-5.4 in ChatGPT for Clinicians) | Physician-written | Margin |
|---|---|---|---|
| Care consult | 51.0 | 42.7 | +8.3 |
| Writing / documentation | 64.1 | 32.1 | +32.0 |
| Medical research | 67.0 | 56.3 | +10.7 |
| Hardest red-team slice (Likert 1–2) | 55.8 | 30.0 | +25.8 |
The widest separation shows up on the adversarial slice. On the hardest red-team examples — the Likert 1–2 difficulty band that physicians deliberately constructed to break models — the tuned product scored 55.8 against 30.0 for physicians and just 26.2 for base GPT-5.4 . The base model and the physicians sit close together at the bottom; the tuned product pulls away. In other words, the adversarial cases are where the specialist tuning earns its margin — and also where the held-out, first-party nature of the evaluation matters most, since red-team difficulty is defined by the same team that graded the results.
One distinction is easy to collapse and worth keeping separate. These 59.0-vs-43.7 figures come from HealthBench Professional, a 525-task clinician-facing suite measuring GPT-5.4 inside a clinician workspace . The consumer-facing "beats doctor-written answers" headline is a different trial: a blind panel comparison across roughly 3,500 responses, reported in the June 18, 2026 health-intelligence post and measuring GPT-5.5 Instant for everyday users . Same framing, two different models, two different exams.
Consumer Answers vs Clinician Consulting: Distinct Trials
Three separate studies now sit behind the "AI matches or beats doctors" framing, and each measures a different system against a different rubric for a different audience. HealthBench, the consumer benchmark, came first; HealthBench Professional reframed the question for clinicians a year later; and the June 18 blind panel is a third, separately reported comparison. Treating them as one result is the easiest way to misread what OpenAI actually demonstrated.
HealthBench, published May 12, 2025, targets lay users. It contains 5,000 realistic consumer health conversations scored against 48,562 physician-written rubric criteria, built with 262 physicians practicing across 60 countries, 49 languages, and 26 specialties . The top reported system, o3, scored 60%, up from 32% for GPT-4o and 16% for GPT-3.5 Turbo, while the 1,000-example HealthBench Hard subset stayed unsaturated at a 32% ceiling . That gap marks where consumer-facing models still fail: context-seeking and worst-case reliability.
HealthBench Professional, submitted to arXiv on April 30, 2026, is a different instrument. Its 525 tasks are clinician-facing — care consult, documentation, and medical research — drawn from 15,079 candidates, with roughly one-third deliberate physician red-teaming and difficult examples enriched about 3.5x . It scores GPT-5.4 inside a clinician workspace, not a lay user typing into free ChatGPT.
The June 18 result is neither. It is a blind panel preference comparison across about 3,500 responses, measuring GPT-5.5 Instant for everyday users . Different product, different rubric, different population. The clinician-facing win belongs to ChatGPT for Clinicians, an institutional tool; the consumer headline belongs to the free default model. As the HealthBench Professional authors caution, the score "is not percent accuracy" and "should not be read as real-world clinical performance rates" .
Conflating the two obscures which system, scored on which rubric, against which population, earned each claim. For developers evaluating these systems, that distinction is the whole signal.
What an Independent Examination Would Find

An independent examination would find no external cross-check. As of June 2026, no outside lab has published a replication of HealthBench or HealthBench Professional scoring, and the design itself blocks exact reproduction: the HealthBench Professional paper, submitted to arXiv on April 30, 2026, notes it uses an internal evaluation implementation and retains a private held-out set, so the headline numbers cannot be regenerated from public materials alone . The benchmark author, physician panel, and tested models all trace to one vendor.
Two specifics deserve scrutiny from any external reviewer:
- The 71% factuality figure. OpenAI reports the rate of responses flagged for at least one possible factuality issue fell 71% over two months, derived from privacy-preserving production monitors at a scale of billions of messages weekly . The threshold definitions and flagging methodology behind that monitor are not disclosed in the public announcement, so the number measures an internal flag rate, not an externally validated error rate.
- The stakes of unaudited claims. Active litigation alleges an earlier model, GPT-4o, gave dangerous medical guidance . When a system reaching the more than 230 million weekly health questioners claims to beat doctors, the gap between a preference signal and a clinical-outcomes trial is not academic.
What would close that gap is third-party replication. Watch for cross-checks from academic health NLP groups or open clinical models such as Meditron and Med-PaLM lineages, which could score the same task families on disclosed rubrics. The regulatory track matters too: FDA guidance treats certain clinical decision-support functions as excluded from the device definition, but software intended for patients or caregivers may still fall under device policy, and AI software as a medical device can require premarket review depending on risk . As OpenAI itself states, ChatGPT Health is not for diagnosis or treatment and is meant to support, not replace, clinician care . Until an independent lab grades the same exam, the result stands as a strong first-party signal, not a settled finding.
Liability and Oversight in Health AI
The legal and regulatory exposure of health AI tracks intended use, and OpenAI's framing is built around that line. By stating ChatGPT Health is meant to support rather than diagnose or treat, OpenAI positions the consumer product outside the strictest device classification — the non-diagnostic claim is the primary lever for any FDA carve-out. U.S. FDA guidance excludes certain clinical decision-support functions from the device definition, but software intended for patients or caregivers may still qualify as a medical device and require premarket review depending on risk and intended use . A "supports, not replaces" label is a positioning choice, not a regulatory exemption.
That distinction matters because the stakes are concrete. There is active litigation alleging an earlier model, GPT-4o, gave dangerous medical guidance . Benchmark wins on physician preference do not resolve liability questions; an answer rated higher than a doctor's by a blind panel can still cause harm in a single case, and rubric scores are not clinical-outcome evidence.
The architecture reflects the compliance gap between consumer and institutional use. ChatGPT Health, announced January 2026, uses separate memories, stores health conversations, apps, and files separately, applies added encryption and isolation, and excludes Health conversations from foundation-model training . These controls are relevant for HIPAA-adjacent contexts but do not by themselves make a free consumer chatbot a covered, audited clinical system.
For that, OpenAI built a distinct tier. OpenAI for Healthcare launched January 8, 2026 with BAA support for HIPAA compliance, PHI controls, evidence retrieval with citations, and policy alignment, naming eight early hospital partners — AdventHealth, Baylor Scott & White Health, Boston Children's Hospital, Cedars-Sinai, HCA Healthcare, Memorial Sloan Kettering, Stanford Medicine Children's Health, and UCSF — with GPT-5.2 models powering the products and PHI kept under organizational control .
The concrete takeaway: there are two products with two risk profiles. The compliance-grade institutional tier sits behind BAAs and PHI controls; the free GPT-5.5 Instant default that "beat" doctors carries only a non-diagnostic disclaimer. Treat the consumer benchmark result as a first-party signal for triage and explanation — not as a substitute for clinical judgment, and not as a settled regulatory or liability position.
Frequently asked questions
What is HealthBench Professional and how does it differ from HealthBench?
Both are OpenAI's own health evaluations, but they target different audiences. The original HealthBench, introduced May 12, 2025, covers consumer-facing care: 5,000 realistic health conversations built with 262 physicians and 48,562 physician-written rubric criteria . HealthBench Professional, posted to arXiv on April 30, 2026, targets clinicians instead: 525 clinician-facing tasks selected from 15,079 candidates, built with 190 physician contributors, with roughly one-third of the set being deliberate physician red-teaming and difficult examples enriched about 3.5x . Same lab, different audience, harder difficulty distribution.
Is the health improvement in GPT-5.5 Instant available on the free tier?
Yes. GPT-5.5 Instant became the default model for free ChatGPT users in May 2026, subject to usage limits, and the health tuning ships at that tier rather than behind a paid plan . That is the consequential change. ChatGPT Health — the dedicated space that connects medical records and wellness apps for personal context — is separate and remains waitlist-gated, with eligibility excluded from the EEA, Switzerland, and the UK, and several record integrations limited to the U.S. .
Why can't the HealthBench Professional results be independently replicated?
OpenAI retains a private held-out set and uses an internal evaluation implementation, so the rubric weights and full task distribution are not published in a form that permits exact reproduction . The arXiv paper makes the methodology partially verifiable, but the headline figures cannot be fully reconstructed from public materials. The lab also built the benchmark, supplied the physician network, and tested its own models — so the superiority claims are first-party signals pending external replication.
Are the 59.0 vs 43.7 scores the same study as the ~3,500 response consumer comparison?
No — they are two distinct trials. The 59.0 vs 43.7 figure comes from HealthBench Professional, where GPT-5.4 running inside ChatGPT for Clinicians scored 59.0 overall against 43.7 for specialty-matched physician-written responses (and 48.1 for base GPT-5.4) . The separately reported consumer study had physicians blindly rate about 3,500 responses and preferred GPT-5.5 Instant over physician-written answers on accuracy, communication, and completeness . Different product, different audience, different rubric.
Does FDA regulation apply to ChatGPT Health?
It depends on intended use. FDA guidance excludes certain clinical decision-support functions from the device definition, but software intended for patients or caregivers may still fall under device policy, and AI software as a medical device can require premarket review depending on risk . OpenAI frames ChatGPT Health as non-diagnostic and meant to support, not replace, clinician care . That positioning is the primary regulatory lever, but it has not been formally tested against the device-definition criteria — and active litigation over an earlier model's medical guidance shows the stakes.