Google SymptomAI tests diagnosis chat on 13,917 real symptom reports
Original: SymptomAI: Towards a conversational AI agent for everyday symptom assessment View original →
The useful test for medical chatbots is not whether they can solve tidy textbook cases. Google Research’s July 22, 2026 SymptomAI study puts the harder question in front: what happens when everyday users describe symptoms with incomplete detail, mixed health literacy, and natural conversation?
The study enrolled 13,917 consenting participants in a national-scale randomized experiment. Each participant interacted with one of five experimental SymptomAI agents built on Gemini Flash 2.0. The agents collected symptom descriptions, asked follow-up questions, and produced a differential diagnosis list. Two weeks later, participants were asked to report any diagnosis they received from a healthcare provider.
Google then ran a clinical expert annotation study. Three board-certified clinicians reviewed the transcripts in blinded fashion and compared differential diagnosis lists from SymptomAI with lists written by clinicians. The standout figure is 53.3%: clinical raters selected SymptomAI as the best first-ranked differential diagnosis more often than the clinician baseline. SymptomAI also more often included the participant-reported provider diagnosis within its top five candidates.
The study’s most practical lesson is about conversation design. Agent-driven history taking outperformed the base condition where the user simply queried a language model. The versions that actively elicited missing information produced stronger diagnostic lists, which suggests that medical AI performance depends heavily on interview structure, not only model scale.
The research also links the conversations to Fitbit biosignals. For participants whose SymptomAI top diagnosis fell into respiratory infection categories, Google observed physiological shifts around the symptom-reporting date across measures such as cardiovascular function, respiration, skin temperature, and sleep. That does not turn the system into a clinical diagnostic product. The paper and blog both stress that the labels are for research analysis only, not confirmed diagnoses. The larger implication is that symptom interviews and wearable data can now be evaluated together at population scale, with safety and clinical validation as the next barrier.
Related Articles
Google Research trained SensorFM on more than one trillion minutes of consented wearable data from five million people. The model beat feature-engineered baselines on 34 of 35 health prediction tasks, pointing to a more general route for wearable health AI.
Frontier AI is becoming both a biosecurity risk surface and a defense tool. Google DeepMind and Isomorphic Labs say they advanced more than 15 partnerships over the past 12 months.
On Feb. 12, 2026, Google announced a major Gemini 3 Deep Think upgrade for science, research, and engineering. The new version is available in the Gemini app for Google AI Ultra subscribers and, for the first time, via early API access for researchers, engineers, and enterprises.