4 must-haves for health execs deploying ambient AI scribes at scale
Physicians love their digital scribes. One recent study showed close to a third of them already using the still-budding technology to capture key content during patient encounters.
The fast rate of uptake may be a good problem to have, but it’s still a problem.
For hospitals and health systems, the rub lies in the technology outpacing validation, transparency and regulatory oversight in siloes all over the enterprise.
A study published March 23 in NPJ Digital Medicine looks at the situation and offers evidence-based insights for healthcare leaders needing to scale ambient AI scribes across diverse and dispersed healthcare settings.
“Ambient AI scribes hold promise in transforming clinical documentation and relieving cognitive and administrative burden as an assistive tool to clinicians,” the authors write. “Yet their success under different care setting hinges not just on technical sophistication but also on ethical design, inclusive evaluation and governance clarity.”
To address these challenges, lead author Joshua Ohde, PhD, of Mayo Clinic and colleagues state that stakeholders “must adopt a systems-level approach grounded in contextual validation, inclusive design and robust governance.”
Among the elements such an approach ought to include, according to the researchers, are these four.
1. Ethical design.
All sorts of ethical concerns and regulatory riddles have come to light as ambient AI scribes have proliferated in clinical environments, the authors point out. Among the examples they cite are problems associated with large-language AI as a whole, not just LLMs in scribes. The authors name model bias and automation bias, hallucinations and potential for misinformation, lack of transparency in training data and legal implications when AI is involved in medical errors.
Within clinical workflows, ethical conundrums extend beyond this to include transparency, privacy, fairness and accountability, Ohde and colleagues note.
“Interestingly, these tools are branded as ‘ambient,’ giving the impression that they are passive and perhaps misguiding as it does not clearly inform patients that their conversations are recorded and saved, often in cloud infrastructure outside the organization’s EHR.”
2. Inclusive development and bias mitigation.
LLM-based systems have been shown to reproduce if not amplify biases present in the training data sets, Ohde and co-authors note. Underrepresented groups may be excluded or misunderstood, they add, if ambient AI scribes are not trained on diverse linguistic patterns, accents and dialects.
“Unfortunately, proprietary models rarely reveal their training and validation data for bias and fairness analysis,” the authors write. “If clinicians are exhibiting automation bias, displaying excessive trust of the tool, the issue could be further compounded.”
Acknowledging that these issues are a particular risk in high-pressure and fast-paced environments, the authors call for cautious evaluation by the adopters themselves.
“Future analyses should also include a qualitative component to elicit the experiences of physicians and patients with the AI scribe based on sociodemographic features,” Ohde and colleagues urge. “Physician–patient communication is often of poorer quality for patients of underrepresented backgrounds, creating potential gaps in outcomes.”
3. Contextual validation.
High-acuity settings may share many similarities, but they tend to be highly diverse across organizations, Ohde and co-authors point out. Variances and peculiarities are not hard to enumerate, they add, by location, physical layout, staffing capabilities, local resources and patient type.
“These differences may impact the effectiveness and safety profile of ambient AI scribes,” the authors write. “Implementation planning, user engagement and post-deployment monitoring are critical to success and risk management under different settings.”
The authors also urge testing of proprietary models across different practice settings. The goal of the testing should be to make sure relative fairness is evident across different patient demographics.
“Currently available, proprietary AI scribes are not easy to adopt or adapt to dynamic user workflows that vary between settings,” Ohde and co-authors write. “Early user engagement to tailor interfaces, workflows and outputs to the realities of clinical practice are necessary. Adaptive training and feedback loops between users and developers will be essential in refining these tools.”
4. Clear and robust governance.
Monitoring in practice should encompass performance, equity and instances of “unsafe acceptance,” Ohde and colleagues maintain. They define the latter terms as “the uncorrected use of a scribe-generated element in pre-identified reduced-reliability contexts later judged incorrect, incomplete or misleading.”
“Continued monitoring of performance drift and unintended adverse events is a critical step in ensuring safety of full -scale implementation; however, current evaluation methods are heavily reliant on human expert evaluation [and] are not scalable,” the authors write.
One potential strategy to circumvent this limitation is to train LLM-based evaluators to assess correctness, relevance, faithfulness, task completion and so on within each high-acuity specialty to augment expert review, Ohde and co-authors offer.
Meanwhile LLM evaluators could “flag notes for manual inspection when issues are identified,” they add. “Such tools need to be tailored to context, undergo regular retraining on new data and be overseen by interdisciplinary governance bodies.”
Ohde and colleagues conclude by underscoring that, with knowledge of current limitations and careful integration, ambient AI scribes can “evolve from passive transcription tools into trusted partners in the delivery of complex care across all care settings.”
The full paper is available for free downloading here.
