Google DeepMind Unveils AI Co-Clinician Research Initiative
Google DeepMind's AI co-clinician recorded zero critical errors in 97 of 98 primary care queries and matched or beat PCPs in 68 of 140 consultation areas in simulated telemedicine trials.

Updated
Why it matters
- WHO predicts a global shortfall of more than 10 million health workers by 2030, the context for DeepMind's AI co-clinician initiative.
- In 98 realistic primary care queries, AI co-clinician recorded zero critical errors in 97 cases, beating two AI systems widely used by physicians under the NOHARM framework.
- In a Harvard and Stanford randomized simulation with 20 scenarios and 10 physician patient-actors, the AI matched or exceeded primary care physicians in 68 of 140 assessed consultation areas, though expert physicians performed better overall.
Google DeepMind has announced an AI co-clinician research initiative, a system designed to work as a collaborative member of the care team that interacts with patients under expert clinical supervision.
The announcement arrives against a global shortage of clinical expertise. The World Health Organization predicts a shortfall of more than 10 million health workers by 2030. DeepMind frames the initiative as a response to that gap: AI that amplifies doctors' expertise rather than replacing it.
"Medicine has always been a team sport, and AI agents can bring more teammates onto the field: extending clinicians' reach while ensuring they retain judgment and control," the DeepMind team writes. The researchers hypothesize the next evolution of healthcare delivery will entail "triadic care" where AI agents help patients through their care journeys under the clinical authority of their physician.
The initiative builds on DeepMind's prior medical AI work, from MedPaLM, which mastered examination-style tests of medical knowledge, to AMIE, which matched physician performance in text-based simulated medical consultations, including in real-world feasibility trial settings.
Evidence evaluation results
For physicians, a tool is useful only if it is trustworthy and factually grounded. Working with academic physicians, DeepMind adapted the "NOHARM" framework to test the system for "errors of commission" (incorrect information) and "errors of omission" (failure to surface critical information).
In head-to-head blind evaluations, physicians consistently preferred AI co-clinician's responses to leading evidence synthesis tools. In an objective analysis of 98 realistic primary care queries, the system recorded zero critical errors in 97 cases, improving over two AI systems widely used by physicians.
The team also evaluated the system on the OpenFDA set of RxQA questions, a benchmark for complex medication knowledge and reasoning. AI co-clinician surpassed other frontier AI systems, especially when questions were posed in the open-ended way they arise in real care.
Real-time multimodal telemedicine
Beyond clinician-facing tools, DeepMind is testing whether AI co-clinician can handle live audio and video in simulated telemedical calls, building on the capabilities of Gemini and Project Astra. Prior studies, including work with Beth Israel Deaconess Medical Center, showed value in AI text chats before a doctor's appointment, but text-only interaction constrains clinical value. "Medicine isn't just text; it requires eyes, ears and a voice," the researchers write.
Working with academic physicians at Harvard and Stanford, the team designed a randomized simulation study with 20 synthetic clinical scenarios and 10 physician "patient-actors." The agent demonstrated capabilities beyond text-only systems, such as guiding patients through complex physical examinations in real time. It successfully corrected a patient's inhaler technique and guided shoulder maneuvers to identify a rotator cuff injury.
The results also set clear limits. Across more than 140 assessed aspects of consultation skill, expert physicians performed better than the AI system overall, particularly in identifying "red flags" and guiding critical physical examinations. The finding suggests these systems are currently best used as supportive tools for practitioners rather than replacements for clinical judgment. Still, AI co-clinician performed at a level comparable to or exceeding primary care physicians in 68 of the 140 assessed areas. Further methodology and results appear in the technical report, "Towards Conversational Medical AI with Eyes, Ears and a Voice."
Safeguards and deployment
The system uses a dual-agent architecture for patient-facing telemedical conversations: a "Planner" module continuously monitors the conversation, verifying that the "Talker" agent stays within safe clinical boundaries. For clinician-facing use, AI co-clinician prioritizes clinical-grade evidence, performing verification and citation checking for retrieval. Physicians constructed the reported evaluations to mirror a range of their real-world evidence needs.
DeepMind is now advancing a phased evaluation approach with academic and research collaborators across healthcare settings in the US, India, Australia, New Zealand, Singapore and the UAE, with plans to expand to mission-aligned healthcare organizations and academic medical centers. The company states its goal is to ensure medical AI is developed and deployed responsibly in line with applicable standards.
The initiative is explicitly research-stage. DeepMind notes the collaborations are not, at this stage, intended for use in the diagnosis, cure, mitigation, treatment, or prevention of disease, or to provide medical advice.
The stakes are substantial. If triadic care proves viable at scale, AI agents could extend scarce clinical expertise across health systems facing a 10-million-worker shortfall — but the simulation results, with physicians still ahead on red flags and physical examination, indicate the near-term role is augmentation, not autonomy.
Original: who.int
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles