Many clinical diagnoses can be derived solely from language-based interviews. Traditionally, clinicians perform these interviews during in-person or remote visits, yet access to such encounters can be limited by financial, geographic, and systemic barriers. Recent language models (LMs) have shown strong differential diagnosis capabilities on curated medical cases, but those evaluations often rely on detailed or synthetic vignettes that may not reflect how real patients describe symptoms in everyday conversation.
To address this gap, the authors carried out an in-situ comparative study of experimental conversational AI agents to investigate how a conversational AI might conduct end-to-end symptom interviews and generate differential diagnostic assessments for research benchmarking. The results are reported in the research paper “SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment.”
Methods
A total of 13,917 consenting participants were enrolled in a randomized national study. Each participant interacted with one of five randomized SymptomAI agents built on Gemini Flash 2.0. The five agents differed in how flexibly they could ask follow-up questions. During each conversation participants described their symptoms; the agent asked follow-ups and concluded with a differential diagnosis (DDx) — a ranked list of plausible diagnoses — plus recommendations for next steps. All labels and diagnoses generated in the study were for research analysis only and did not constitute confirmed clinical diagnoses.
Two weeks after their interaction, participants were surveyed about any diagnoses they received from a healthcare provider following the conversation. To evaluate SymptomAI’s assessments, a panel of three board-certified clinicians performed a blinded clinical annotation study: they reviewed the conversation transcripts, produced their own DDx, and then blindly ranked the DDx lists produced by SymptomAI and by the clinicians.
In addition, the study compared SymptomAI’s categorizations against wearable biosignals. Consenting participants provided daily Fitbit biometric data for up to 30 days prior to their SymptomAI conversation to explore physiological trends leading up to symptom reporting.
Key results
-
Clinician preference for SymptomAI DDx: In more than 50% of reviewed cases, clinicians preferred the DDx generated by SymptomAI over those provided by other clinicians. This suggests SymptomAI’s suggestions matched or exceeded peer clinicians’ assessments at least as often as clinicians’ assessments matched each other.
-
Higher top-5 accuracy: The study compared whether the diagnosis that participants later reported receiving from a healthcare provider appeared in the top five items of each DDx list. Clinicians judged the SymptomAI-generated DDx to contain the eventual provider diagnosis more often than the DDx lists provided by other clinicians.
-
Eliciting more information improves performance: Participants were randomized into five arms using different prompting strategies: Dynamic Live and Dynamic Final (agents had unrestricted ability to ask follow-up questions), Fixed Canonical and Flexible Canonical (agents asked from a standard set of history-taking questions taught in medical school), and Base (an unprompted LM representing user-driven chat). All agent-driven prompting strategies significantly outperformed the Base condition, demonstrating the value of active follow-up questioning for improving differential diagnostic accuracy.
-
Greatest gains on low-confidence cases: SymptomAI’s advantage over clinical baselines was largest for cases where clinicians expressed lower confidence in their own DDx.
-
Diagnosis correlates with biosignals: For cases where SymptomAI’s top-1 candidate was classified as an infectious respiratory illness (excluding non-infectious respiratory conditions such as allergic rhinitis or chronic obstructive pulmonary disease), the authors observed clear shifts in wearable biometrics in the days leading up to symptom reporting. These shifts included changes in cardiovascular metrics, respiration, skin temperature, and sleep quality, with peaks aligned to the symptom report date. The alignment provides observational physiological evidence that corresponds to the AI-derived symptom classifications.
-
Utility for population-scale research: The authors note that automated, accurate symptom assessment systems like SymptomAI could enable large-scale labeling of clinical-quality diagnoses for population analyses of physiological data, a task currently constrained by the cost of clinical labels.
Limitations
The authors outline several limitations:
-
Differential diagnosis is an inherently ambiguous task and diagnoses can evolve longitudinally. A symptom assessment reflects a moment in time; because of the scale of the deployment, the timing and frequency of symptom reports could not be controlled. Some participants may have reported symptoms very early or, conversely, reported well-established chronic indicators.
-
In the clinician evaluation, reviewers saw only static chat transcripts and could not ask their own follow-up questions; clinicians might have collected different information had they conducted the interview themselves. Conversational AI may also miss non-verbal or external signals such as body language, visual assessment, or prior medical records, and it lacks any pre-existing rapport a clinician might have with a patient.
-
All study-generated diagnoses, labels, and disease associations are AI-derived for research analysis only and do not constitute clinical diagnoses.
Conclusion
SymptomAI is an investigational conversational AI intended for research on real-world patient interviews and symptom assessment. In a large randomized sample (n=13,917), SymptomAI produced DDx lists that clinical experts often preferred and rated as accurate more frequently than DDx lists from other clinicians. Additionally, SymptomAI’s infectious diagnoses correlated with pre-conversation wearable biosignal changes, suggesting potential for combining conversational symptom assessment with passive physiologic data in population-scale health research.
Acknowledgements
The work reflects equal contributions from Joe Breda, Jake Sunshine, and Daniel McDuff. The authors thank co-authors and collaborators from Google Research and Google DeepMind for their contributions.



