Results on unseen patients.

We scored 12,431 held-out patients every hour, then 27,484 more in three outside databases without retraining. All of it is retrospective.

48 hours before kidney failure.

Scroll to move through the hours. The patient is scored every hour and flagged the first time their risk crosses the review line.

Hour
−48 h
Status
Below review threshold
Flag
Not flagged

Illustrative, not a real patient.

Strongest on liver and kidney.

EndpointHorizonAUROCAUPRC
Circulatory failure8 h0.9360.656
Hyperglycaemia8 h0.9340.584
Sepsis8 h0.8450.546
Kidney failure48 h0.9620.896
Liver failure48 h0.9820.943
Macro average0.9320.725

Outside databases, no retraining.

Weaker in children, who are outside our current scope.

Macro AUROC across the five endpoints

Better than bedside scores in adults.

Against NEWS2, partial SOFA and lactate, reconstructed hour by hour. It doesn't beat them in children, where lactate does just as well.

AUROC for any deterioration event

Zentus compositeBedside score

Full tables.

Model facts
Intended use
An hourly estimate of each adult ICU patient's risk of five deterioration events, for silent validation only.
Intended users
ICU clinical and research teams reviewing logged output after a pilot.
Scope
Adults; silent validation, with no output shown at the bedside.
Outcomes and horizons
Circulatory failure, hyperglycaemia and sepsis, 8 hours ahead; kidney and liver failure, 48 hours ahead.
Validation
12,431 held-out patients from MIMIC-III, MIMIC-IV and eICU; 27,484 external patients from NWICU, Zigong and PICDB, without retraining.
Discrimination
AUROC 0.845 to 0.982 by endpoint on the internal test set; macro 0.932.
Calibration
Two of five endpoints are overconfident in the top decile and will be recalibrated before any alerting.
Regulatory status
Research software; not cleared or approved by any regulator; not for clinical use.
Evidence version
Retrospective evidence packet dated 17 May 2026; peer review pending.
Patients and stays in each evaluation set
CohortSplitPatientsICU stays
MIMIC-III, MIMIC-IV, eICUInternal test12,43117,411
NWICUExternal17,34920,260
ZigongExternal2,2552,255
PICDB paediatricExternal7,8808,378
External totalExternal27,48430,893

No patient appears in more than one split.

Each endpoint in each external database
DatabaseEndpointAUROCAUPRCPositive rate
NWICUCirculatory failure0.9480.1240.11%
NWICUHyperglycaemia0.9610.6083.59%
NWICUKidney failure0.9400.4160.97%
NWICULiver failure0.9600.2370.10%
NWICUSepsis0.8640.0450.47%
ZigongCirculatory failure0.8990.4020.53%
ZigongHyperglycaemia0.9350.3774.70%
ZigongKidney failure0.8510.3200.95%
ZigongLiver failure0.8940.3831.74%
ZigongSepsis0.6900.1396.35%
PICDBCirculatory failure0.9060.2370.65%
PICDBHyperglycaemia0.8670.1701.12%
PICDBKidney failure0.7740.0720.08%
PICDBLiver failure0.7320.0140.77%
PICDBSepsis0.7110.0280.89%

Rolling zero-shot results. Events are rare here (0.1% to 6.4% of patient-hours), so AUPRC values are unstable; sepsis transfers least well.

Composite against bedside scores: AUPRC
CohortCompositeNEWS2 equiv.Partial SOFALactate
Internal test0.8290.3960.5980.651
NWICU0.1310.0210.0260.054
Zigong0.3110.0790.1670.220
PICDB0.0300.0170.0320.074

Oxygen supplementation and mental status are not fully harmonised across cohorts. Patient-hours: internal 1,176,991; NWICU 649,553; Zigong 16,834; PICDB 171,258.

Calibration, top decile
EndpointPredictedObserved
Liver failure99%99%
Kidney failure96%95%
Sepsis65%62%
Circulatory failure74%58%
Hyperglycaemia70%48%

Circulatory failure and hyperglycaemia run high. We'll recalibrate both before any alerting.

Not measured yet
  • These results measure prediction. The effect of acting on it is a later study.
  • Confidence intervals and subgroup results (age, sex, unit) are the next analysis.
The internal threshold used at other sites
CohortSensitivitySpecificityPPVFlag rate
Internal test80%86%74%36%
NWICU57%94%10%6.5%
Zigong34%93%32%9%
PICDB50%71%3%30%

Composite threshold 0.17, set for 80% sensitivity on the internal validation split. This is why a pilot logs against thresholds fixed in advance, and any alerting threshold comes from the site's own silent log.

Flag rate at two operating points
EndpointHigh sensitivity
flags / patient-day
High sensitivity
stays flagged
Conservative
flags / patient-day
Conservative
stays flagged
Circulatory failure0.733.5%0.066.3%
Hyperglycaemia2.034.7%0.169.6%
Sepsis1.554.3%0.087.8%
Kidney failure0.614.8%0.051.6%
Liver failure0.212.5%0.022.1%

High sensitivity: set for 80% sensitivity on the validation split. Conservative: flags 1% of patient-hours; in this retrospective run it caught no true kidney or liver events, so it is not suitable for those endpoints as set. Internal rolling test set, with a cooldown between repeat flags.

Discuss a silent pilot in your unit.

Validation on a Malaysian hospital's own data is next.