A Gemini-powered medical AI combined sight, sound, and clinical reasoning to match or outperform physicians on key measures in simulated video consultations, while revealing where human connection and subtle perception still matter.
*AMIE guiding a patient actor through a physical examination (pronator drift test) over a video call. Figure 1 from the paper, 'Towards Expert-level Medical AI for Real-time Video Consultations', enhanced using generative AI. *
*Important notice: ** arXiv publishes preliminary scientific reports that are not peer-reviewed and, therefore, should not be regarded as conclusive, guide clinical practice/health-related behavior, or treated as established information.
A recent study published on the arXiv
Challenges and Opportunities in Audio-Visual Clinical AI
Video-based consultations are the dominant modality for remote clinical care, enabling clinicians to observe nonverbal cues, perform limited physical examinations, and establish rapport. In contrast, most demonstrations of medical AI have focused on instant messaging (IM) interfaces, which excel at structured reasoning but fail to capture the complexity of live patient interactions.
Early studies have shown that extending medical AI to audio-visual modalities is feasible, paving the way for more naturalistic and effective remote care. Nonetheless, substantial challenges persist, including the need for accurate, real-time perception of audio-visual cues, adaptability across diverse patient presentations, seamless, low-latency dialogue management, and the integration of sophisticated clinical reasoning. Achieving clinician-level performance in these domains remains an open and critical objective for next-generation AI clinical agents.
AMIE Development and Evaluation for Clinical Use
Google researchers developed AMIE, a multi-agent system for real-time, video-based clinical consultations, integrating three specialized agents: the Talker Agent (for rapid, fluent patient responses), the Planner Agent (for managing clinical goals), and the Perception Agent (for interpreting audio-visual data). Built on Gemini 3 Flash and 3.1 Pro, these agents work asynchronously to generate clinically appropriate, low-latency dialogue.
Development was guided by an automated evaluation framework including targeted single-turn tests and simulated multi-turn conversations to assess end-to-end clinical and conversational performance. The evaluation taxonomy spans a wide range of clinical audio-visual cues and physical maneuvers relevant to telehealth, distinguishing cues by virtual feasibility and actability by patient actors. Evaluations use scenario-specific rubrics and large language model (LLM)-based auto-raters.
For clinical evaluation, a randomized multi-arm Objective Structured Clinical Examination (OSCE) study compared AMIE (in both video and text configurations) with 10 US board-certified primary care physicians (PCPs) across 100 simulated patient scenarios. Consultations were performed with 15 professional standardized patient actors and scored using both general and case-specific rubrics by a panel of 20 PCP clinical evaluators.
Assessment criteria included diagnostic accuracy, clinical reasoning, management, communication, and perceptual acuity. Additional qualitative data from patient actors and evaluators further assessed user experience.
AMIE Matches or Exceeds PCPs in Simulated Real-Time Video Settings
AMIE outperformed or matched PCPs in simulated real-time video consultations. Clinical evaluators rated AMIE Video as equal to or better than PCPs across all five case-specific clinical domains and gave it a higher overall score: 83% compared to 68%.
Patient actors preferred AMIE Video for assessing and explaining conditions, though there were no significant differences in empathy or addressing concerns. AMIE Video achieved higher diagnostic accuracy, with its top-ranked differential diagnosis matching the prespecified reference diagnosis in 91% of scenarios compared to 77% for PCPs. This difference narrowed and was not statistically significant when the top three differential diagnoses were considered (98% vs. 90%).
AMIE also excelled in clinical reasoning, scoring 90% versus 76%, and in treatment planning, with 79% compared to 67%. History-taking scores were also higher for AMIE, and its use of live video and guidance for patient-assisted physical examinations was substantially stronger than that of PCPs. AMIE was particularly strong in leveraging video for live observation and guiding exams. Communication scores favored AMIE, though patient actors showed a non-significant preference for PCPs in rapport and partnership building.
AMIE Interaction Example showcasing activities of Talker, Perception and Planner agents.
When comparing video and text chat interfaces, patient actors rated the video version significantly higher in terms of communication effectiveness, ease of use, and understanding concerns. Video offered the greatest improvements in perception and examination skills. Diagnostic and management performance were similar across both modalities, while individual patient-actor ratings of agent behavior were also largely similar, although they generally favored video.
Qualitative insights indicated that video consultations enabled direct demonstration of symptoms and consistent expressions of empathy, and that AMIE was described as thorough and unhurried. PCPs, however, were favored for natural conversational flow. Some limitations for AMIE Video were noted, including occasional missed visual cues and technical challenges.
Automated evaluations showed that using specialized agents, especially Perception and Planner, greatly boosted AMIE’s performance. The complete system outperformed the Talker-only version overall in both single-turn and multi-turn clinical scenarios, excelling in visual assessments and clinical reasoning. In the multi-turn automated simulations, however, visual examination findings were conveyed through verbal stage directions rather than continuous live video.
While AMIE delivered strong results in most physical domains, it still struggled with subtle perceptual tasks, such as detecting nystagmus, tremor, or interpreting mood. Crucially, the asynchronous agent design enabled higher automated-evaluation performance while reducing mean turn latency from 21.4 to 2.6 seconds compared with the earlier sequential architecture.
After each live encounter, AMIE and the physicians also completed structured post-encounter questionnaires covering differential diagnoses and management. For this less latency-constrained step, AMIE used multiple drafts and synthesis to improve the quality of its recommendations.
Next Steps for Clinical Video AI
AMIE's asynchronous, multi-agent approach demonstrates strong potential for advancing AI-assisted clinical video consultations. However, limitations remain, including challenges with fine anatomical precision, affective nuance, and the detection of high-frequency movements, as well as the reliance on simulated patients rather than real-world clinical encounters. The system also relies on discrete turn-taking, preventing the natural overlapping dialogue and visually triggered interjections possible during human consultations.
The comparison also differed from typical video consultations because PCPs were required to keep their cameras off to match AMIE's lack of a visual avatar, which the authors noted may have affected rapport and empathy ratings. Full blinding was not possible because communication styles and voices differed between AMIE and physicians.
In the future, researchers should focus on addressing these limitations and conducting extensive validation within real clinical workflows to ensure both safety and effectiveness. As the technology matures, AMIE and similar systems can expand access to high-quality care and enhance telemedicine capabilities.
*Important notice: ** arXiv publishes preliminary scientific reports that are not peer-reviewed and, therefore, should not be regarded as conclusive, guide clinical practice/health-related behavior, or treated as established information.