Google's medical AI, AMIE, has passed a milestone that headlines have flattened into a single number. In a prospective study run at Beth Israel Deaconess Medical Center (BIDMC) and published in The Lancet, the conversational diagnostic system interviewed 100 real patients — with physicians supervising — and its suggestions were compared against the diagnosis each patient eventually received.
That comparison is where the reporting gets slippery. AMIE listed the eventual diagnosis among its top seven possibilities in 90% of cases, among its top three in 75%, and as its first choice in 56% of 98 evaluable patients. "90% diagnostic accuracy" is therefore the top-seven figure, not a hit rate for the AI's leading answer. Clinicians who have reviewed the study have been blunt about the gap: an assistant that is broadly right about the neighbourhood is useful, but it is not the same as one that gets the answer right first time.
The distinction matters because the study is the first prospective evaluation of AMIE in a real clinic rather than a simulated or retrospective benchmark, making it the most consequential evidence so far about whether a doctor-facing AI can help in ordinary practice. AMIE's role here was preparation: gathering history and generating a differential before the clinician saw the patient, which is precisely the kind of low-risk, high-leverage task hospitals are considering first.
Google Research framed its conclusions carefully. The study suggests AMIE "has potential to improve clinicians' diagnostic reasoning and accuracy in challenging cases, meriting further real-world evaluation" — a claim about assisting doctors, not replacing them. The caveats are substantial: one site, a hundred patients, physician oversight throughout, and a single snapshot rather than a deployed clinical workflow.
The wider context is a crowded race to put AI in front of clinicians. Google's own earlier work demonstrated AMIE in live video consultations matching doctors in simulated settings; the same companies now routinely publish studies that mix benchmark results with real-world pilots. For hospitals weighing adoption, the practical question raised by this Lancet paper is not whether a model can name the right disease in a list, but whether it changes what a physician does, misses less, and does no harm — questions only deployment-scale evidence can answer.
Sources
- thelancet.comConversational diagnostic artificial intelligence in ambulatory care (The Lancet)
- sites.research.googleAMIE (Google Research)
- arise-ai.orgARISE, BIDMC and Google AMIE publish prospective clinical AI study
- explainx.aiGoogle AMIE Lancet Study: 90% Hit Rate, Not Doctor-Level (explainx.ai)
- superpowerdaily.comGoogle's medical AI helped doctors prepare for visits in small Lancet study




