Head-to-head evaluation of adult inpatient deterioration predictive models enables clinical leaders to de-risk AI deployments

For health system leaders struggling with high rates of inpatient mortality and ICU transfers, evaluating AI/ML models that predict deterioration on local patient data can build clinician confidence and reveal operational tradeoffs for clinical workflows.

Failing to proactively respond during the early stages of inpatient deterioration events has implications for patient mortality, readmission rates, and hospital length of stay. Early identification provides care teams an opportunity to intervene before a patient’s condition becomes critical, preventing avoidable deaths and keeping patients from being rushed into the ICU. Today, clinical leaders can choose from a wide range of predictive AI/ML models and rules-based systems to detect patients at-risk of inpatient deterioration events – but performance and results vary widely across algorithms and clinical settings.

A multistate health system compared the performance of four different models across 118,063 historical patient encounters leveraging the Vega Health Platform. The evaluation equipped leaders with insights into each model’s predictive performance, clinical and operational metrics like alert burden, and optimal strategies for implementation in clinical workflows.

Results from the retrospective cohort indicated that the Duke Deterioration Index model outperformed both a widely used vendor deterioration index model and the National Early Warning Score (NEWS) 2 Scoring System as a predictive tool in this health system’s clinical environment. The evaluation also equipped health system leadership with evidence that combining a rules-based score (e.g., NEWS 2) with a predictive model (e.g., Duke Deterioration Index) could improve accuracy, help identify more deterioration events before they occur, and provide more clinically interpretable patient insights to bedside care teams than a predictive model alone.

Using AI to Predict Deterioration

Chart-based scoring systems like NEWS 2 can give clinicians a way to evaluate patient stability and identify patients at-risk of deterioration. Health systems, such as Duke Health, Kaiser Permanente, and the University of Michigan, have also developed and deployed more sophisticated AI/ML algorithms to predict deterioration risk and support proactive clinician decision-making for at-risk patients. EHR vendors also make predictive deterioration models widely available5 to health systems.

With so many options available, clinical leaders need a way to identify the optimal model and approach for their care teams and patient populations. Retrospective model evaluations provide a powerful tool to establish local performance and generate evidence for health system leaders to select and configure AI for clinical use.

Optimizing Outcomes with Combined Algorithm Deployment

The predictive ability of an AI model is only part of the picture. When health systems deploy a model, they must also understand how it will affect clinical workflows. Operational performance measures, such as alert burden and the number of missed events, provide insights on how the models may perform in real-world clinical settings. Low sensitivity can lead to missed interventions and poor outcomes, while a high volume of false alerts can erode clinician trust, contribute to “alert fatigue,” and divert valuable resources from the bedside.

At Duke Health, the predictive Duke Deterioration Index model was deployed with phenotype flags to support clinician workflows (see Duke Health Case Study for additional insights). For systems that want to prioritize explainability for bedside nurses, this approach also has the ability to improve sensitivity and increase clinician confidence.

A retrospective evaluation gave health system leadership the opportunity to evaluate the combined deployment of NEWS 2 and the Duke Deterioration Index model before making an implementation decision. The combined approach captured more deterioration events than either model individually. Over the two year evaluation, NEWS 2 alone predicted 64.7% of deterioration events. By adding Duke Deterioration Index alerts, the system may have captured an additional 19.8% of total deterioration events.

Identifying Risk & Bias

Retrospective evaluations can also provide health system leaders with insights into where a model should not be deployed. When models were evaluated for each population subgroup – age group, biological sex, preferred language, and race – performance was typically within expected bounds. However, subgroup analyses revealed performance variation for patients in Labor & Delivery, which informed decisions about which units to include in the scope of an implementation and where additional evaluation would be required.

Make AI Deployment Decisions with Confidence

As predictive models become more commonplace in clinical settings, health system leaders should not risk failed pilots, alarm fatigue, and poor patient outcomes. A local retrospective model evaluation provides both confidence in model performance and insights into operational considerations before deployment – better positioning health systems to improve patient outcomes and deliver measurable results.

Note to the Reader: The Duke Deterioration Index model and phenotypes were included in the head-to-head comparison through a distribution partnership between Duke Health and Vega Health. The widely used vendor deterioration index model and NEWS 2 Scoring System were included to help inform a health system partner’s AI deployment. For other health systems that may be interested in a similar model evaluation, there will likely be variability in the results based on the system’s patient population and workflows, which is why rigorous, site-specific benchmarking matters.

Download Full Report