The Duke Institute for Health Innovation (DIHI) has spent the better part of the last decade building AI models for some of Duke Health’s thorniest problems. One of the top focus areas for developers around the country has been identifying patients who are at risk of a rapid deterioration in their condition because early intervention can significantly impact the ultimate outcome.

The Duke Deterioration Index, a predictive AI model, was built to flag patients at risk of an unplanned ICU transfer or in-hospital death hours before it happens.

A follow-on to Sepsis Watch

The project began in 2019 as a follow-on to Sepsis Watch, DIHI's sepsis prediction and management solution. Dr. Cara O'Brien, a clinical lead on Sepsis Watch, came back through DIHI's Request for Application (RFA) process with a new problem. Duke already had an internal decompensation model, but it ran only every eight hours and wasn't catching patients in time to act.

The eventual Duke Deterioration Index model was deliberately simpler than Sepsis Watch. Sepsis Watch used a recurrent neural network that also had to forecast where each input would be at regular intervals, which made it expensive to keep running. For deterioration, the team chose a gradient-boosted tree model instead.

Predicting the call before it happens

The model estimates each patient's risk of deteriorating in the next 12 hours, defined as an unplanned ICU transfer or in-hospital death, and refreshes that risk every 15 to 20 minutes. The goal, in Suresh's words, is to predict a call to the rapid response team well before it's made.

That creates an immediate challenge: if a patient crosses the risk threshold, they are likely to remain above it 20 minutes later. Sending the same alert to the same clinicians with each refresh quickly creates alarm fatigue, making it more likely that important EHR alerts will be ignored.

So, Duke built a “snooze button.” Once a patient triggers an alert, repeat alerts are held for a set window, about six hours in Duke's current setup. The window roughly matches the rhythm of care: most labs and vitals are rechecked every four hours or so, so a new alert after the snooze reflects new information. Suresh noted a second benefit: fewer repeat alerts means a higher share of alerts are true positives.

The right alert to the right person

Because the model looks 12 hours ahead, a flagged patient often looks fine when a nurse walks in. Duke's answer was to make chart review the first step, not an afterthought. At Duke University Hospital, a medium, high or critical risk page goes to the Patient Response Team charge nurse, who reviews the chart and then does a quick bedside assessment.

If there's no concern, the patient is handed back to the bedside nurse. If there is, the rapid response team puts a care plan in place, from new orders and assessments to procedures like intubation.

That workflow varies by campus and by unit. At Duke University Hospital, alerts route to the Patient Response Team charge nurse; at Duke Regional Hospital, they go to the Early Nurse Intervention Team; and at Duke Raleigh Hospital, they go to the Rounding Nurse Team. The differences reflect how each site staffs, rounds, and escalates concerns on individual units. Suresh emphasized that localization has to happen at the unit level, not just the hospital level. The design should start with two practical clinical questions: who needs to be notified, and who is accountable for acting on the alert.

From a risk score to a clinical reason

For most of its life, the model was a black box: clinicians got a risk score and nothing else. That worked well enough to route alerts to the rapid response team, but it left nurses to figure out why a patient was flagged.

About a year ago, Duke began pairing the model with its Deterioration Phenotypes. These are six rules-based flags defined by Duke physicians: hypotension, end organ dysfunction, hypoperfusion, vasoactive medications, respiratory decline and respiratory intervention. Each is a yes-or-no check against clear clinically validated thresholds, so a nurse can see exactly what's happening.

At Duke, the risk model and the Deterioration Phenotypes initially ran as separate tools for four to five years. Duke now combines them so an alert fires only when a patient's risk score crosses the threshold and the patient meets two or more phenotype criteria. Suresh said this change significantly improved positive predictive value, while giving clinicians a clearer starting point for chart review and bedside assessment.

The first test outside Duke

Until now, the Duke Deterioration Index had never been validated outside Duke Health. Its results had been shared internally and at the Machine Learning for Healthcare (MLHC) conference, never in a peer-reviewed publication. Evaluating the EHR vendor's deterioration model had not been necessary at Duke, because Duke already had its own models in use.

Vega Health tested the model on two years of data from a multistate health system: 118,063 adult inpatient encounters across five hospitals. We compared it with a widely used vendor deterioration index and NEWS 2, and recreated Duke's combined deployment with the phenotypes. What we found tracked closely with Duke's experience:

  • Better discrimination. The Duke model had an AUROC of 0.812, compared with 0.748 for the vendor model and 0.714 for NEWS 2 (p < 0.001 for both comparisons).
  • More events for the same alerts. At the vendor model's high-alert rate, the Duke model caught 55% more deterioration events.
  • The snooze works. A 4-hour snooze after the first alert cut Duke alerts per patient-day by 62% without reducing the share of events captured.
  • Phenotypes raise confidence. Among high-risk Duke alerts, 5.5% of patients went on to deteriorate. That rose to 8.8% with two phenotypes present and 17.2% with three or more.

The evaluation also showed where not to use the model. Performance in Labor & Delivery (L&D) was well outside expected bounds, which isn't surprising since the model excludes those patients from training. Further analysis and refinement is needed before using it in L&D.

These results come from one health system, and results elsewhere will vary with patient populations, data quality and workflows. That's exactly why site-specific evaluation and localization matters.

What comes next

For Suresh, the deterioration work sits inside a broader DIHI effort to catch patients before they decompensate, alongside Sepsis Watch and models that identify patients at risk of dying so care teams can start goals-of-care conversations. Seeing how the Deterioration Index performs at another health system also feeds back into how DIHI builds its next generation of clinical decision support models, so they're easier to develop, deploy and maintain elsewhere with similar performance.

The lesson from Duke's years of running the model is that a strong performance score is only the start. Thoughtful unit-level localization of adapting alerts to local workflows, using snoozes to reduce alert fatigue, and pairing model output with actionable context is what gives nurses the information they need to act at the bedside.