Who: leaders from academia and healthcare policy gathered at Johns Hopkins University’s Carey Business School, including Ziad Obermeyer of UC Berkeley, Zachary Lipton of Carnegie Mellon University, and physician educator Dr. Emily Boss of Johns Hopkins. What: they discussed how health systems can learn from peer organizations when adopting artificial intelligence in clinical care. When: the event took place at the end of May. Where: the discussion was hosted at the Johns Hopkins Bloomberg Center in Washington, D.C. Why: because real-world AI deployments increasingly depend on rigorous evaluation, fairness checks, and workflow integration—not just model performance in initial testing.
Context: Why peer lessons matter in healthcare AI
Artificial intelligence in healthcare has moved beyond research pilots, with vendors and providers rolling out tools for tasks such as imaging support, clinical documentation, risk prediction, and operations planning. Yet organizations still report uneven results when models transition from controlled studies to busy clinical environments.
Several widely documented problems drive that gap. AI systems can produce biased outputs when training data reflect historical inequities, and they can lose accuracy as populations, care pathways, and device calibration change over time. The U.S. Food and Drug Administration (FDA) maintains public listings of AI/ML-enabled medical devices, underscoring that evaluation happens in regulatory frameworks—but it does not eliminate post-deployment performance risk (FDA, AI/ML-enabled medical devices database).
A core lesson from earlier algorithm failures is that performance metrics alone may not reveal who benefits and who is harmed. In a landmark paper published in Science, researchers including Obermeyer found that an algorithm used to identify patients who would benefit from more care relied on health-care costs as a proxy for need, which correlated differently across groups and led to systematic under-identification for Black patients (Obermeyer et al., Science, 2019).
That background sets the stage for the event’s focus: health systems can accelerate safer deployment by comparing approaches—how they evaluate models, monitor drift, and govern data—rather than treating each rollout as a one-off experiment.
Main themes: What leaders said about borrowing from peers
The panel framed
Frequently Asked Questions
What does it actually mean for health systems to “learn from peers” when adopting AI?
It means comparing approaches that go beyond choosing a model: how organizations evaluate it, run fairness checks, and test it in real clinical workflows. Instead of treating each rollout as a one-off experiment, peers can share playbooks for validation design, monitoring practices, and governance decisions that reduce risk during and after deployment.
Why do performance metrics from initial testing often fail once AI reaches everyday clinical environments?
Because clinical reality changes. Models can lose accuracy as populations shift, care pathways evolve, or devices and calibration drift. Metrics captured in controlled studies may not reflect these operational conditions. Peer discussions focus on readiness for workflow integration and ongoing evaluation, not just early model performance.
If the FDA lists AI/ML medical devices publicly, doesn’t that guarantee ongoing safety?
No. FDA evaluation helps establish baseline assessment within regulatory processes, but it does not remove post-deployment performance risk. Real-world effects like data drift, changing patient mix, and evolving clinical practice can alter model behavior over time. That’s why health systems still need monitoring and governance after go-live.
How can an algorithm appear accurate overall while still causing harm to specific groups?
Because overall metrics can hide unequal impact. A well-known example from Science (2019) showed an algorithm that used healthcare costs as a proxy for need, correlating differently across groups. That led to systematic under-identification for Black patients, illustrating why fairness assessment must be part of evaluation—not an afterthought.
What should health systems monitor after deployment to catch problems early?
Peers emphasized monitoring drift and re-evaluating performance in the context of real workflows. That includes tracking whether outputs remain consistent as populations and care pathways change, verifying continued accuracy, and reviewing fairness signals over time. Pairing monitoring with governance ensures issues are acted on rather than only reported.

