Research

AI-generated text

Real-time, expert-level AI for video medical consultations demonstrated in simulated study

Researchers present AMIE (Video), a real-time, multi-agent AI system that perceives audio-visual clinical cues, guides virtual physical exams, and performs diagnostic reasoning.

Real-time, expert-level AI for video medical consultations demonstrated in simulated study

Google Research and Google DeepMind researchers present AMIE (Video), a research AI system designed for real‑time, audio‑visual medical consultations. The system is intended to overcome limitations of text‑only interfaces by perceiving nonverbal visual and auditory cues, guiding virtual physical examinations, and performing diagnostic reasoning synchronously.

Why this matters

In clinical practice, nonverbal signals—gait, visible discomfort, breathing patterns, and responses to examination maneuvers—contribute substantially to diagnosis and patient communication. Text-only interfaces discard these signals and require patients to translate complex physical findings into written descriptions, which can lose diagnostic information and disadvantage patients with limited digital or health literacy.

Architecture and operation of AMIE (Video)

Built on Gemini and Project Astra, AMIE (Video) runs synchronous video consultations using an asynchronous multi‑agent architecture that splits tasks across three parallel, specialized agents:

  • Talker agent: the patient-facing module that maintains low‑latency spoken interaction and the natural flow of conversation while incorporating guidance from other agents.
  • Planner agent: operates in the background to continuously refine clinical reasoning, update differential diagnoses and management plans, and identify information gaps.
  • Perception agent: continuously analyzes audio and video streams to detect clinically relevant nonverbal cues (for example, signs of respiratory distress or visible physical findings) and situates those observations within the ongoing dialogue.

This decoupling lets AMIE (Video) preserve conversational responsiveness while performing deeper diagnostic reasoning and audio‑visual perception that would otherwise introduce prohibitive delays if done serially.

Development guided by automated evaluation

To characterize perceptual and reasoning capabilities at scale, the team derived a taxonomy of clinical audio‑visual competencies from the medical literature covering visual cues, auditory signals, and physical exam maneuvers. They built an automated evaluation suite that combines:

  • single‑turn audio‑visual assessments focused on specific perception and reasoning tasks (e.g., identifying anatomical laterality or recognizing respiratory distress), and
  • multi‑turn simulated audio consultations that inject visual cues as textual descriptions into the simulation (for example, a Parkinson’s scenario where an AI patient simulator verbalizes “[holding up paper to camera showing cramped, tiny script]”).

These complementary evaluations supported rapid iteration, detailed characterization of capabilities, and identification of failure modes before human evaluation.

Randomized video OSCE study

The team evaluated end‑to‑end clinical competence in a large randomized Objective Structured Clinical Examination (OSCE) conducted via synchronous video consultations. Key study details:

  • 100 clinical scenarios covering five body systems: cardiopulmonary, abdominal, HEENT (head/eyes/ears/nose/throat), neurological/psychiatric, and musculoskeletal.
  • 15 trained patient actors carried out 300 standardized consultations across three study arms.
  • Study arms: AMIE (Video); AMIE (Text), a text‑only baseline; and PCP (Video), in which physicians consulted via the same video interface. The paper reports a group of 30 board‑certified primary care physicians (PCPs) participating in the research and indicates that ten board‑certified PCPs provided consultations in the PCP (Video) arm.
  • An independent panel of 20 experienced primary care physicians evaluated all consultations using established clinical rubrics and case‑specific scoring criteria.

Main findings

  • Expert‑level clinical performance: Evaluators rated AMIE (Video) comparable to PCPs on core competencies, including thoroughness of history taking, diagnostic accuracy, management appropriateness, and communication quality. AMIE (Video) matched or exceeded the text‑only AMIE (Text) on these dimensions.

  • Strength in physical observation and examination: AMIE (Video) was rated significantly higher on average than both PCPs and AMIE (Text) for eliciting physical signs and proactively guiding patient actors through virtual examination maneuvers. This advantage appeared in case‑specific perception and exam rubric scores.

  • Patient actor preference: Patient actors strongly preferred the synchronous video interface over text chat and rated AMIE (Video) favorably for empathy, rapport, and confidence in care relative to both PCPs and AMIE (Text).

Limitations and responsible development

The authors emphasize several limitations:

  • The study used professional patient actors in simulated encounters rather than real patients presenting their own health conditions. Actors cannot fully replicate the complexity and unpredictability of real clinical encounters, and the scenarios were limited to conditions that can be authentically portrayed.
  • Automated evaluations identified occasional perceptual and reasoning errors, and the system displayed intermittent technical issues that can disrupt conversational naturalness.
  • Project Astra is prototypical; some technical considerations will need system‑level development beyond the medical application itself.

Because of these limitations, validation with real patients and real clinical conditions is necessary before assessing real‑world utility.

Next steps

The work shows that moving from text‑based to audio‑visual clinical AI is achievable at expert‑level quality in simulated settings. Important next steps include validating findings with real patients, expanding to clinical presentations that cannot be enacted by actors, and establishing robust safety and oversight frameworks.

The authors note prior and ongoing translational efforts: a real‑world feasibility study with Beth Israel Deaconess Medical Center provided initial evidence for the safety and utility of text‑based AMIE, and a nationwide randomized study with Included Health is further evaluating AI in real virtual care settings. These efforts aim to inform how audio‑visual capabilities could be responsibly integrated into clinical practice.

Acknowledgements

The research is joint work across many teams at Google Research and Google DeepMind. The authors acknowledge numerous co‑authors, including Mahvish Nagda, Jihyeon Lee, Matthew Thompson, CJ Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, and others.


(This article summarizes research findings from simulated studies; further clinical trials with real patients are required to determine real‑world effectiveness and safety.)