Google's medical conversational AI just crossed a line that matters for anyone building clinical agents: from naming a disease to managing one over time. In a June 2026 Nature paper, AMIE's chronic-management plans were rated at least as good as those of primary care physicians across simulated, multi-visit cases.
What the longitudinal chronic management extension does differently
The change is a shift from one-shot diagnosis to management reasoning across multiple visits. Earlier versions of AMIE (Articulate Medical Intelligence Explorer) produced a ranked differential diagnosis from a single consultation — no follow-up planning, no tracking of treatment response between encounters. The extended system instead generates investigations, prescriptions, and per-visit follow-up protocols for conditions like congestive heart failure and type-2 diabetes, and updates them as symptoms evolve and test results arrive .
That capability was evaluated at a scope no prior published AMIE study covered: 100 chronic-disease scenarios spanning five specialties, each played out over three consultation visits . The work began as the preprint "Towards Conversational AI for Disease Management" (arXiv:2503.06074, submitted 8 March 2025) and was reported as a peer-reviewed Nature publication in mid-June 2026, led by Mike Schaekermann and colleagues at Google Research and DeepMind .
To grade plans that unfold over time, the team introduced a new scoring instrument — MXEKF, a management-reasoning rubric for longitudinal plan quality. Across 51 comparison sets, the median tie rate was 50%; where evaluators expressed a preference, AMIE's median win rate was 42% versus 8% for PCPs . The takeaway for builders: this is a longitudinal evaluation protocol, not a single-turn benchmark — and the rest of this analysis unpacks how the simulation was run and where the margins actually held up .
How the chronic-disease blinded simulation was structured

The evaluation was a randomized, blinded virtual OSCE (objective structured clinical examination) run entirely over text chat — not a clinical deployment, and with no EHR integration, audio, video, labs, or imaging . Both AMIE and the human comparators answered through the same text interface, with management plans stripped of any authorship signal before grading. That symmetry is the methodological core: evaluators could not tell which plan came from a model and which from a physician.
The comparator pool was 21 board-certified primary care physicians from India and North America, with 21 trained patient actors playing the cases and 10 specialist physician evaluators scoring the output, blinded to authorship, across 15 evaluation axes . The case set covered 100 scenarios spread evenly across five specialties — cardiology, pulmonology, OB/GYN/urology, gastroenterology, and neurology, 20 each .
Each scenario ran across three visits, and the cases were built to stress longitudinal reasoning rather than one-shot diagnosis: evolving symptoms, treatment responses, incoming test and imaging results, deliberate information inconsistencies, and multimorbidity. To probe portability across jurisdictions, scenarios were prepared in both Canada and India .
Critically for a fair comparison, AMIE and the PCPs drew on the same reference material — a 627-document corpus of 527 NICE Guidance documents plus 100 BMJ Best Practice documents, roughly 10.5 million tokens in total . Neither side held a guideline-access advantage; the variable under test was how each used it.
"Management plans were graded by specialists blinded to authorship, with AMIE and primary care physicians communicating through the same text interface" — Mike Schaekermann and colleagues, describing the study design (source: Nature, 2026).
The design is worth dwelling on because it bounds the claim. This is a controlled simulation with actors and a UK-oriented guideline base (NICE/BMJ), not a study of real patients in live care . With that scaffolding established, the next question is where the measured margins actually held up.
Treatment and investigation precision: margins and significance
On the metrics that mattered most, AMIE's edge over PCPs was precision, not volume. In the blinded simulation, treatment precision — the share of recommended treatments judged appropriate — ran 94% vs 67% at visit 1, 90% vs 70% at visit 2, and 91% vs 70% at visit 3, every gap statistically significant . The advantage was widest at the first consultation and narrowed slightly across follow-ups, which Google attributed to AMIE's relative lead being strongest at the initial encounter .
Quick Answer: In the simulated OSCE, AMIE's treatment precision hit 94% vs 67% for PCPs at the first visit and stayed near 90% across all three visits — every margin statistically significant . Its strength was fewer unnecessary tests and treatments, not broader recall.
Investigation precision told a compounding story. The two were tied at 91% at visit 1, but AMIE pulled ahead at follow-up: 99% vs 84% (p=0.009) and 100% vs 88% (p<0.001) . In other words, the gap opened as cases evolved across visits rather than appearing all at once.
| Metric | Visit 1 (AMIE / PCP) | Visit 2 | Visit 3 |
|---|---|---|---|
| Treatment precision | 94% / 67% | 90% / 70% | 91% / 70% |
| Investigation precision | 91% / 91% | 99% / 84% | 100% / 88% |
Beyond precision, AMIE led on plan-level judgment at the first visit: overall plan appropriateness was 88% vs 74% (p=0.019), appropriate follow-up recommendations 100% vs 98% (p<0.001), and appropriate treatment at the first follow-up 85% vs 71% (p=0.014) .
The pattern across all 15 evaluation axes and three visits was non-inferiority — AMIE never graded worse than PCPs — with the standout gains clustered on precision rather than recall . Read carefully, that distinction matters for builders: the system's win was ordering fewer unnecessary tests and treatments, not catching more conditions overall. Precision-without-recall-loss is the harder bar in chronic care, where over-investigation drives cost and patient burden. Independent reporting framed these results, alongside the separate MIRA paper, as evidence that conversational clinical AI is approaching physician-level performance on management tasks — though, as the next section shows, the precision edge tracks closely with how well each side grounded its plan in guidelines.
Guideline grounding across all consultation rounds

AMIE grounded its plans in clinical guidelines more consistently than primary care physicians at every consultation round, and that grounding tracked its precision edge directly. At visit 1, specialists rated AMIE higher on selecting the applicable guideline — 92% versus 76% for PCPs (p=0.016) — and the advantage held at both follow-up visits . In other words, the system was better at picking the right rulebook before it started prescribing, which is the upstream step that makes a management plan defensible.
Treatment recommendations followed the same pattern. AMIE's plans were more guideline-aligned across all three encounters — 89% vs 75%, 91% vs 76%, and 93% vs 81% — with the gap widening slightly by visit 3 . The system also showed its work: a far higher share of its recommendations carried an explicit guideline reference (98% vs 86%, 100% vs 87%, 100% vs 88%) . For a developer, that traceability is the auditable part — every recommendation effectively ships with a citation.
| Guideline metric | Visit 1 (AMIE/PCP) | Visit 2 | Visit 3 |
|---|---|---|---|
| Applicable guideline selected | 92% / 76% | — | — |
| Treatment guideline alignment | 89% / 75% | 91% / 76% | 93% / 81% |
| Explicit guideline reference cited | 98% / 86% | 100% / 87% | 100% / 88% |
The mechanism behind this is the Mx Agent's retrieval budget. It allocated roughly 256,000 context tokens per case to external knowledge — averaging about six applicable guideline documents filtered from a 627-document corpus (527 NICE Guidance plus 100 BMJ Best Practice documents, ~10.5 million tokens) . Both AMIE and the PCPs could access the same corpus, drawn from NICE Guidance and BMJ Best Practice; the difference was how reliably each side retrieved and referenced it.
RxQA: a new scoring challenge for medication knowledge
RxQA is a 600-question multiple-choice benchmark for medication reasoning, and on its hardest items AMIE beat primary care physicians while neither side cleared 75% on the easier ones . The questions are derived from two live national formulary sources — OpenFDA and the British National Formulary — then revised by board-certified pharmacists, covering indications, contraindications, dosages, and drug interactions . That sourcing is the point: unlike single-turn academic sets such as MedQA or MedMCQA, RxQA is pulled from formulary references that track current prescribing guidance, which makes it a usable standalone evaluation instrument rather than a static exam dump.
The headline numbers are sobering for anyone assuming pharmacology is a solved benchmark. On lower-difficulty open-book questions, AMIE scored 73.8% against the PCPs' 67.4% — a gap that did not reach significance (p=0.071), and a ceiling that surprised on the low side for both, given full reference access . The model's edge only became clear on the harder, pharmacist-rated items:
- Higher-difficulty, closed-book: AMIE 50.6% vs PCPs 41.5% (p=0.013)
- Higher-difficulty, open-book: AMIE 57.9% vs PCPs 47.8% (p<0.001)
The pattern is consistent: the separation holds on the questions that demand deeper retrieval and reasoning, not on the routine ones. For builders, that argues against treating medication accuracy as binary — a model can look competent on easy lookups and still leave roughly half of the hard interaction-and-dosage questions wrong.
"RxQA fills a gap left by single-turn benchmarks, giving a concrete way to test medication reasoning against current formulary guidance," note the study authors describing the instrument's purpose (source: Nature, 2026-06).
Paired with the multi-visit OSCE protocol from the same study, RxQA gives clinical-AI developers a reusable yardstick for the part of management reasoning that is easiest to get plausibly — but dangerously — wrong.
Dialogue front end, deliberative back end: the clinical AI split

The architecture that produced those scores splits clinical conversation from clinical reasoning into two agents on the Gemini model family. A user-facing Dialogue Agent — built on Gemini 1.5 Flash, replacing the earlier PaLM-2 base — converses with the patient in near-real time, manages conversational state across the encounter, and maintains an empathetic tone at low latency . A separate Management Reasoning ("Mx") Agent does the slow, deliberative work: synthesizing multi-visit dialogue plus retrieved clinical guidance into a management plan under a target response time of about one minute per encounter .
That one-minute target is the tell: the Mx path is explicitly not built for synchronous chat. Splitting it off lets the deliberative component cross-reference authoritative sources without stalling the conversation. The Mx Agent filters a 627-document corpus — 527 NICE Guidance documents and 100 BMJ Best Practice documents, roughly 10.5 million tokens — and allocates about 256,000 context tokens to the most relevant subset, corresponding on average to about six applicable guideline documents per case .
For builders, the pattern transfers beyond medicine. The design choices map cleanly onto a general recipe for compliance-grounded conversational systems:
- Route latency-sensitive conversation to a fast model. Gemini 1.5 Flash keeps the front end responsive and holds session state.
- Route compliance-critical, document-grounded reasoning to a long-context model on an async path. The Mx Agent spends its token budget on retrieval and synthesis, not turn-by-turn responsiveness.
- Share state between encounters. The two components persist patient context across visits, which is what makes longitudinal management — not just one-off Q&A — possible.
The separation also isolates the part most likely to fail audit. Keeping guideline retrieval and plan synthesis in a dedicated agent makes that reasoning inspectable and groundable, rather than entangled with conversational generation (source: Google Research, 2026-06).
Simulation limits, caveats, and prospective feasibility
Every number in this study comes from simulation, not the clinic — which bounds how far the result can travel. The entire evaluation ran as text-chat consultations with trained patient actors, with no electronic health record (EHR) integration and no audio, video, lab, or imaging inputs . Real patients arrive with comorbidities, language barriers, and incomplete histories that scripted OSCE scenarios do not fully reproduce, so physician-level performance on a virtual exam does not yet imply physician-level performance in care.
Two further constraints matter for anyone weighing deployment:
- UK-only guideline grounding. The Mx Agent's corpus was drawn exclusively from NICE Guidance and BMJ Best Practice . Localization to other jurisdictions, formularies, and clinical settings is untested — the system's strongest signal, guideline alignment, is calibrated to one country's rulebook (NICE, BNF).
- No real outcomes measured. The blinded OSCE rated plan quality, not morbidity, mortality, adherence, hospitalization cost, or long-term chronic-disease control . A 94% treatment-precision score is a rater judgment, not a patient outcome.
The closest thing to live evidence sits in a separate domain. A March 2026 single-arm feasibility preprint put AMIE into real urgent-care history-taking: 100 adults completed chats up to five days before their visits, no human safety supervisor had to intervene, and AMIE's differential included the final diagnosis in 90% of cases with 75% top-3 accuracy . But blinded reviewers found PCPs beat AMIE on practicality (p=0.003) and cost-effectiveness (p=0.004) — and that study was diagnosis, not chronic management. Live longitudinal management has not been attempted.
The authors are explicit about the gap. The paper states AMIE is not ready for clinical care and calls for prospective feasibility studies with ethical and safety oversight before any integration . Google says next steps include exploring clinical-setting integration and a nationwide randomized study in real-world virtual care, with no deployment timeline announced (source: Google Research, 2026-06).
The takeaway for builders: treat this as a validated benchmark protocol and architecture pattern, not a clinical claim. The multi-visit OSCE design, the RxQA medication benchmark, and the Dialogue+Mx split are reusable today; the precision margins are a simulated signal that earns a prospective trial, not a green light. Until EHR-integrated, multi-jurisdiction, outcome-measured studies exist, "matches PCPs" means "matches PCPs on a blinded text exam" — a meaningful result, and a deliberately narrow one.
Frequently asked questions
How does AMIE's longitudinal management capability differ from its original diagnostic mode?
The original AMIE handled a single consultation and produced a differential diagnosis — a ranked list of probable conditions from one dialogue. The extended system reasons about management over time, tracking disease progression and treatment response across multiple visits to produce visit-specific investigations, prescriptions, and follow-up protocols rather than a one-off diagnosis . It was evaluated on 100 three-visit scenarios and scored with a new management-reasoning rubric, MXEKF; across 51 comparison sets the median tie rate was 50%, and where a preference was expressed AMIE's median win rate was 42% versus 8% for PCPs .
What is RxQA and why should clinical AI developers pay attention to it?
RxQA is a 600-question medication-reasoning benchmark built from two national formulary sources — OpenFDA and the British National Formulary (BNF) — and revised by board-certified pharmacists, covering indications, contraindications, dosages, and interactions . Unlike single-turn benchmarks such as MedQA, it derives from live formulary data (OpenFDA, BNF), so it can track current prescribing guidance instead of a static snapshot. It also stays hard: on higher-difficulty pharmacist-rated questions AMIE scored 50.6% vs 41.5% closed-book (p=0.013) and 57.9% vs 47.8% open-book (p<0.001) .
Why can't the OSCE simulation stand in for clinical safety validation?
Because it was a blinded text-chat exam, not care. All 100 cases used trained patient actors over text — no EHR integration, no real patients, and no multimodal input (audio, video, labs, or imaging) . Guidelines came from a single UK-oriented set (NICE/BMJ), so other jurisdictions and formularies are untested, and the study measured no morbidity, mortality, adherence, or cost outcomes . The authors state explicitly that AMIE is not ready for clinical care and call for prospective feasibility studies with ethical and safety oversight .
Can I reuse the Dialogue + Mx split in my own clinical AI build?
Yes — the architecture is the most portable part of the work. It separates a responsive conversational front end (the Dialogue Agent, built on Gemini 1.5 Flash for low-latency synchronous chat) from a deliberative, document-grounded back end (the Management Reasoning "Mx" Agent, a long-context model on an asynchronous path targeting about a one-minute response) . The Mx Agent retrieved from a 627-document corpus (527 NICE plus 100 BMJ documents, ~10.5M tokens) and allocated roughly 256,000 context tokens per case . The pattern fits any system where compliance-critical reasoning against a curated corpus must coexist with a responsive chat interface.
How does this relate to the MIRA paper in the same Nature window?
MIRA (Medical Intelligence for Reasoning and Action), from Jakob Kather's group, targets diagnosis and treatment decision support in emergency settings, while AMIE targets longitudinal chronic-disease management in primary care . Both appeared in Nature in mid-June 2026 and were framed jointly as evidence that conversational clinical AI is approaching physician-level performance in controlled simulation . The shared caveat matters: neither is a deployed or cleared product, and both results come from simulated rather than real-world clinical conditions.