Results on unseen patients.
We scored 12,431 held-out patients every hour, then 27,484 more in three outside databases without retraining. All of it is retrospective.
48 hours before kidney failure.
Scroll to move through the hours. The patient is scored every hour and flagged the first time their risk crosses the review line.
- Hour
- −48 h
- Status
- Below review threshold
- Flag
- Not flagged
Illustrative, not a real patient.
Strongest on liver and kidney.
| Endpoint | Horizon | AUROC | AUPRC |
|---|---|---|---|
| Circulatory failure | 8 h | 0.936 | 0.656 |
| Hyperglycaemia | 8 h | 0.934 | 0.584 |
| Sepsis | 8 h | 0.845 | 0.546 |
| Kidney failure | 48 h | 0.962 | 0.896 |
| Liver failure | 48 h | 0.982 | 0.943 |
| Macro average | 0.932 | 0.725 |
Outside databases, no retraining.
Weaker in children, who are outside our current scope.
Macro AUROC across the five endpoints
Better than bedside scores in adults.
Against NEWS2, partial SOFA and lactate, reconstructed hour by hour. It doesn't beat them in children, where lactate does just as well.
AUROC for any deterioration event
Full tables.
Model facts
- Intended use
- An hourly estimate of each adult ICU patient's risk of five deterioration events, for silent validation only.
- Intended users
- ICU clinical and research teams reviewing logged output after a pilot.
- Scope
- Adults; silent validation, with no output shown at the bedside.
- Outcomes and horizons
- Circulatory failure, hyperglycaemia and sepsis, 8 hours ahead; kidney and liver failure, 48 hours ahead.
- Validation
- 12,431 held-out patients from MIMIC-III, MIMIC-IV and eICU; 27,484 external patients from NWICU, Zigong and PICDB, without retraining.
- Discrimination
- AUROC 0.845 to 0.982 by endpoint on the internal test set; macro 0.932.
- Calibration
- Two of five endpoints are overconfident in the top decile and will be recalibrated before any alerting.
- Regulatory status
- Research software; not cleared or approved by any regulator; not for clinical use.
- Evidence version
- Retrospective evidence packet dated 17 May 2026; peer review pending.
Patients and stays in each evaluation set
| Cohort | Split | Patients | ICU stays |
|---|---|---|---|
| MIMIC-III, MIMIC-IV, eICU | Internal test | 12,431 | 17,411 |
| NWICU | External | 17,349 | 20,260 |
| Zigong | External | 2,255 | 2,255 |
| PICDB paediatric | External | 7,880 | 8,378 |
| External total | External | 27,484 | 30,893 |
No patient appears in more than one split.
Each endpoint in each external database
| Database | Endpoint | AUROC | AUPRC | Positive rate |
|---|---|---|---|---|
| NWICU | Circulatory failure | 0.948 | 0.124 | 0.11% |
| NWICU | Hyperglycaemia | 0.961 | 0.608 | 3.59% |
| NWICU | Kidney failure | 0.940 | 0.416 | 0.97% |
| NWICU | Liver failure | 0.960 | 0.237 | 0.10% |
| NWICU | Sepsis | 0.864 | 0.045 | 0.47% |
| Zigong | Circulatory failure | 0.899 | 0.402 | 0.53% |
| Zigong | Hyperglycaemia | 0.935 | 0.377 | 4.70% |
| Zigong | Kidney failure | 0.851 | 0.320 | 0.95% |
| Zigong | Liver failure | 0.894 | 0.383 | 1.74% |
| Zigong | Sepsis | 0.690 | 0.139 | 6.35% |
| PICDB | Circulatory failure | 0.906 | 0.237 | 0.65% |
| PICDB | Hyperglycaemia | 0.867 | 0.170 | 1.12% |
| PICDB | Kidney failure | 0.774 | 0.072 | 0.08% |
| PICDB | Liver failure | 0.732 | 0.014 | 0.77% |
| PICDB | Sepsis | 0.711 | 0.028 | 0.89% |
Rolling zero-shot results. Events are rare here (0.1% to 6.4% of patient-hours), so AUPRC values are unstable; sepsis transfers least well.
Composite against bedside scores: AUPRC
| Cohort | Composite | NEWS2 equiv. | Partial SOFA | Lactate |
|---|---|---|---|---|
| Internal test | 0.829 | 0.396 | 0.598 | 0.651 |
| NWICU | 0.131 | 0.021 | 0.026 | 0.054 |
| Zigong | 0.311 | 0.079 | 0.167 | 0.220 |
| PICDB | 0.030 | 0.017 | 0.032 | 0.074 |
Oxygen supplementation and mental status are not fully harmonised across cohorts. Patient-hours: internal 1,176,991; NWICU 649,553; Zigong 16,834; PICDB 171,258.
Calibration, top decile
| Endpoint | Predicted | Observed |
|---|---|---|
| Liver failure | 99% | 99% |
| Kidney failure | 96% | 95% |
| Sepsis | 65% | 62% |
| Circulatory failure | 74% | 58% |
| Hyperglycaemia | 70% | 48% |
Circulatory failure and hyperglycaemia run high. We'll recalibrate both before any alerting.
Not measured yet
- These results measure prediction. The effect of acting on it is a later study.
- Confidence intervals and subgroup results (age, sex, unit) are the next analysis.
The internal threshold used at other sites
| Cohort | Sensitivity | Specificity | PPV | Flag rate |
|---|---|---|---|---|
| Internal test | 80% | 86% | 74% | 36% |
| NWICU | 57% | 94% | 10% | 6.5% |
| Zigong | 34% | 93% | 32% | 9% |
| PICDB | 50% | 71% | 3% | 30% |
Composite threshold 0.17, set for 80% sensitivity on the internal validation split. This is why a pilot logs against thresholds fixed in advance, and any alerting threshold comes from the site's own silent log.
Flag rate at two operating points
| Endpoint | High sensitivity flags / patient-day | High sensitivity stays flagged | Conservative flags / patient-day | Conservative stays flagged |
|---|---|---|---|---|
| Circulatory failure | 0.7 | 33.5% | 0.06 | 6.3% |
| Hyperglycaemia | 2.0 | 34.7% | 0.16 | 9.6% |
| Sepsis | 1.5 | 54.3% | 0.08 | 7.8% |
| Kidney failure | 0.6 | 14.8% | 0.05 | 1.6% |
| Liver failure | 0.2 | 12.5% | 0.02 | 2.1% |
High sensitivity: set for 80% sensitivity on the validation split. Conservative: flags 1% of patient-hours; in this retrospective run it caught no true kidney or liver events, so it is not suitable for those endpoints as set. Internal rolling test set, with a cooldown between repeat flags.
Discuss a silent pilot in your unit.
Validation on a Malaysian hospital's own data is next.