Monday, July 20, 2026

News

Google's Medical AI AMIE Matches Doctors in Managing Chronic Disease

ResearchPatryk Raba1

A study published in Nature shows that Google's AMIE medical AI performed at least as well as 21 primary care physicians in managing patients across multiple visits, and outperformed them on some tasks. The authors caution that the system is not yet ready for real clinical use.

Contents
  1. From Diagnosis to Ongoing Treatment
  2. How the Study Was Conducted
  3. Testing Drug Knowledge
  4. Caveats and Next Steps

Google published the results of a study on its medical AI system called AMIE in the journal Nature, testing it for the first time not on a single diagnosis but on managing a patient over a series of follow-up visits. In a blind, randomized experiment, the system matched 21 primary care physicians in overall treatment reasoning, and in several categories, such as treatment plan accuracy and adherence to clinical guidelines, it scored significantly higher.

From Diagnosis to Ongoing Treatment

AMIE, short for Articulate Medical Intelligence Explorer, is a system built on Google's Gemini models, developed by Google Research and Google DeepMind over several years. Earlier versions focused on a single diagnostic conversation simulating a visit to a primary care doctor. The new study tests something different: the system's ability to stay with a patient over time, as symptoms change, test results arrive gradually, and treatment needs adjusting.

In practice, this means AMIE has to remember findings from previous visits, track the response to previously prescribed medication, and update the treatment plan according to clinical guidelines that are themselves regularly revised. According to Google, it is this stage, not the diagnosis itself, that generates the most administrative work and errors in clinical practice.

Giving a diagnosis is the first step in treating a patient. Once a diagnosis is established, the challenge becomes managing a health condition over time - tracking symptoms across multiple appointments, parsing guidelines as they're updated and fine-tuning medications - Mike Schaekermann, Research Lead, Google Research

How the Study Was Conducted

Researchers ran a randomized, blind study modeled on the OSCE exam format, a standard method for assessing the clinical competence of medical students. AMIE and the 21 primary care physicians conducted text-based conversations with trained actors playing patients, across 100 scenarios each covering three consecutive visits. Symptoms, test results, and treatment response changed between visits, mirroring real long-term care.

Independent specialists, who did not know whether they were evaluating a plan prepared by AMIE or by a doctor, rated the system's care plans as no worse than the doctors' plans in terms of overall treatment reasoning. In categories such as the relevance of recommended diagnostic tests, medication dosing accuracy, and explicit references to specific clinical guidelines, AMIE scored higher than the group of doctors.

Testing Drug Knowledge

A separate part of the study was the RxQA pharmacological knowledge test, made up of 600 questions based on official FDA databases and the UK's British National Formulary. The questions covered indications, contraindications, dosing, and drug interactions. According to Google, the system performed particularly well on higher-difficulty questions, outperforming the doctors' results.

AMIE runs on an architecture made up of two cooperating components. One, a dialogue agent, carries the ongoing conversation with the patient, gathers information, and builds rapport. The other, a reasoning agent, analyzes the collected clinical data against hundreds of pages of medical guidelines and generates a structured treatment plan, using the long context window of the Gemini models to process a patient's history across multiple visits at once.

Caveats and Next Steps

The study's authors explicitly state that AMIE is not ready for use in real clinical care. The experiment consisted solely of simulated, text-based consultations with actors rather than real patients, and no actual health outcomes were measured. The scenarios were designed to evaluate the system, not to reflect the typically chaotic day-to-day reality of primary care. The system also remains prone to confabulation, generating false or made-up answers, which carries obvious risks in a medical context.

Google says it plans further, more rigorous testing under conditions closer to real-world practice before any version of the tool could be used in actual clinical settings. The company points to the potential for systems like AMIE to help ease global doctor shortages and inequalities in access to healthcare, while stressing the need for further, quantitative research into exactly how the system arrives at its recommendations.

For Poland's healthcare system, where patients can wait months for a specialist appointment and family doctors see dozens of patients a day, findings like these carry practical weight even though real-world deployment remains distant. The publication of the AMIE results fits a broader trend of testing AI as a tool to support doctors with repetitive, time-consuming parts of care, such as updating treatment plans in line with changing guidelines, rather than replacing them in clinical decisions.

Share: