Research prototype — not for clinical use. MSc dissertation project; results are internal validation and one external transportability check, not a clinical claim.
Predicting heart
failure outcomes,
honestly.
Five models, four outcome horizons, 2,008 patients — and a full leakage audit that cut the headline AUROC from 0.819 to 0.710 before I believed a single digit of it. Then I retrained inside MIMIC-IV to test whether the neurological signal was real.
0
Zhang patients
Heart-failure clinical records
0
Final predictors
After the leakage audit
0
MIMIC-IV admissions
External transport set
0
AUROC from GCS
MIMIC-IV, 28-day, grouped split
The live prototype
Now runs the real models.
The deployed Streamlit app serves the audited pipeline — same 143 predictors, same thresholds, with a research-only framing.
Launch the live HF-RISK tool in its own workspace
Real inference
Same pipeline as the dissertation.
Live SHAP
Per-patient explanations on demand.
Research-safe
Framed as prototype, never advice.
The audit
Higher AUROC isn't better if it's not honest.
An earlier pass reached 0.819 at six months. The audit found features that leak the future — and the honest number is 0.710 internal, 0.599 external.
Before audit
0.819
After audit
0.710
Removed as leakage
- Unnamed: 0
- inpatient.number
- dischargeDay
- outcome.during.hospitalization
- readmission & emergency-return timing proxies
Every removed feature was available only after the outcome was already decided — perfect hindsight, zero clinical value.
Results
The numbers that survive scrutiny.
Held-out Zhang cohort, four horizons, two model families.
0.781
28-day mortality
Random Forest
0.839
3-month mortality
Random Forest
0.710
6-month mortality
XGBoost
0.650
6-month readmission
XGBoost
2026 extension · GCS
Does GCS add anything beyond an ICU stay?
GCS topped the Zhang SHAP ranking, so I retrained inside MIMIC-IV — 42,990 admissions from 18,890 patients, grouped by patient — and tested the components against a real ICU-stay control.
28-day mortality
0.8834 0.8972
+0.0138 AUROC · 95% CI 0.0090–0.0194
6-month mortality
0.8266 0.8352
+0.0086 AUROC · 95% CI 0.0054–0.0120
ICU subgroup, 6-month descriptive
0.8557 0.8766
+0.0208 AUROC · 95% CI 0.0141–0.0274
| Outcome | Model | Features | Test AUROC | Δ vs baseline (95% CI) |
|---|---|---|---|---|
| 28-day | Expanded baseline | 34 | 0.8834 | Reference |
| 28-day | + GCS components | 37 | 0.8972 | +0.0138 (0.0090–0.0194) |
| 28-day | + ICU-stay flag | 35 | 0.8915 | +0.0080 (0.0037–0.0125) |
| 28-day | + ICU flag + GCS | 38 | 0.9000 | +0.0166 (0.0119–0.0220) |
| 6-month | Expanded baseline | 34 | 0.8266 | Reference |
| 6-month | + GCS components | 37 | 0.8352 | +0.0086 (0.0054–0.0120) |
| 6-month | + ICU-stay flag | 35 | 0.8280 | +0.0014 (−0.0012–0.0040) |
| 6-month | + ICU flag + GCS | 38 | 0.8370 | +0.0104 (0.0073–0.0136) |
| ICU only, 6-month | Baseline | 34 | 0.8557 | Reference |
| ICU only, 6-month | + GCS components | 37 | 0.8766 | +0.0208 (0.0141–0.0274) |
Shaded rows: a confidence interval crossing zero (no detectable improvement at this sample size), and the ICU-subgroup result, which uses a single split rather than five-fold cross-validation.
Why the split matters
8,006 of the 18,890 patients have more than one heart-failure admission (one has 59). A row-level split would scatter the same person across train and test and inflate every score. Every number above comes from a grouped 80/20 split on subject_id with zero patient overlap asserted before scoring.
GCS is left missing, never imputed
GCS is charted for 12,083 of 42,990 admissions (28.1%) — the rest were never assessed, so the field is left empty and XGBoost handles the NaN natively. gcs_total is excluded as a feature: for intubated patients it reads 15 regardless of the components. Treating an untested patient as a normal score would fabricate a clinical observation.
The ICU control
The ICU-stay flag comes from the icustays table, independently of whether GCS was charted. At six months the flag alone is a null result (+0.0014, CI crossing zero) while GCS still adds +0.0104 on top of it — so the gain is not simply “was this patient sick enough for intensive care”.
Process
Five weeks, four tracks.
Develop, audit, explain, validate — in parallel, not in series.
- Step 01
Develop
Five candidate models, consistent splits, no peeking.
- Step 02
Audit
Feature-by-feature leakage review against the data dictionary.
- Step 03
Explain
SHAP reduction to 143 predictors that carry the signal.
- Step 04
Validate
Held-out cohort, MIMIC-IV transfer, calibration, DCA, subgroups.
Feb 20–27
Foundations & literature
Feb 28 – Mar 4
Exploratory analysis
Mar 4–8
Preprocessing & features
Mar 9–14
Modelling & comparison
Mar 14–18
Explainability & tool
Mar 18–20
Outreach & dissertation
Explainability
What the final model actually pays attention to.
Mean |SHAP| across the held-out cohort — the cardiorenal axis, in descending order.
GCS (Glasgow Coma Scale)
100%
Moderate–severe CKD
88%
Mitral valve AMS
78%
LV end-diastolic diameter
72%
Liver disease
66%
CHF history
60%
Eye opening (GCS sub-score)
54%
Reduced EF flag
48%
Basophil ratio
42%
Creatine kinase
36%
Why dischargeDay had to go
Day-of-discharge correlates with when the record was written, not with the patient. It inflated every horizon. Removing it cost 0.109 AUROC at six months — and bought a model that means what it says.
A note on GCS
In the Zhang cohort GCS tops the chart partly because it is 15 for 97.2% of patients — its signal is mostly “is it ever low?”. The MIMIC-IV extension answers that properly: leave the score missing instead of imputing a normal one, add an independent ICU-stay flag as a control, and the components still earn a small, consistent increment.
External validation
Transportability to MIMIC-IV.
The model was trained on Zhang cohort records and applied to 42,990 MIMIC-IV admissions with no retraining — a transportability check, not a clinical claim.
Under true domain shift the 28-day AUROC falls from 0.781 to 0.524 — a 0.257 drop that leaves it barely above chance. Reported as a transportability check, not a clinical claim. Retraining on local data is the honest next step, which is exactly what the 2026 GCS extension does.
Evaluation
Calibration, DCA, and subgroup checks.
A model is only as good as its worst-calibrated decile. Brier 0.027, ECE 0.023.
Glossary
Terms, plainly.
BNP
Range 2.7–5,000 pg/mL in this cohort; danger > 400; mean ≈ 1,280.
LVEF
Healthy 55–70%; reduced < 40%; missing for 68% — imputed with MICE.
NYHA
Class I–IV heart-failure severity; 52% Class III, 31% Class IV here.
AUROC
Chicco & Jurman benchmark ≈ 0.73 for this dataset; final 0.710 at six months.
SHAP
Game-theory attributions: how much each feature moved this prediction.
MICE
Multivariate imputation by chained equations — used on 42 variables.
References
Five sources, no padding.
- Zhang et al. (2021). Heart failure clinical records. Scientific Data 8:46. doi:10.1038/s41597-021-00835-9
- Goldberger et al. (2000). PhysioBank, PhysioToolkit, and PhysioNet. Circulation 101(23), e215–e220.
- Johnson et al. MIMIC-IV. PhysioNet.
- Chicco & Jurman (2020). The advantages of the Matthews correlation coefficient. BMC Med Inform Decis Mak 20:16 — baseline AUC ≈ 0.73.
- Lundberg et al. (2020). SHAP. Nature Machine Intelligence.