Research prototype — not for clinical use. MSc dissertation project; results are internal validation and one external transportability check, not a clinical claim.

MSc Data Analytics Heart Failure ML 2026

Predicting heart
failure outcomes,
honestly.

Five models, four outcome horizons, 2,008 patients — and a full leakage audit that cut the headline AUROC from 0.819 to 0.710 before I believed a single digit of it. Then I retrained inside MIMIC-IV to test whether the neurological signal was real.

Author
Nikunj Prajapati
Institution
London Met.
Course
MSc Data Analytics
Cohort
Zhang / PhysioNet
External set
MIMIC-IV (42,990)
Updated
Oct 2026 · GCS extension

0

Zhang patients

Heart-failure clinical records

0

Final predictors

After the leakage audit

0

MIMIC-IV admissions

External transport set

0

AUROC from GCS

MIMIC-IV, 28-day, grouped split

The live prototype

Now runs the real models.

The deployed Streamlit app serves the audited pipeline — same 143 predictors, same thresholds, with a research-only framing.

hf-risk — render service — live inference Checking…

Launch the live HF-RISK tool in its own workspace

Real inference

Same pipeline as the dissertation.

Live SHAP

Per-patient explanations on demand.

Research-safe

Framed as prototype, never advice.

Open the live HF-RISK tool → View source on GitHub

The audit

Higher AUROC isn't better if it's not honest.

An earlier pass reached 0.819 at six months. The audit found features that leak the future — and the honest number is 0.710 internal, 0.599 external.

Before audit

0.819

After audit

0.710

Outcome horizonBeforeAfterΔModel
28-day mortality 0.893 0.781 −0.112 Random Forest
3-month mortality 0.919 0.839 −0.080 Random Forest
6-month mortality (primary) 0.819 0.710 −0.109 XGBoost
6-month readmission 0.648 0.650 +0.002 XGBoost

Removed as leakage

  • Unnamed: 0
  • inpatient.number
  • dischargeDay
  • outcome.during.hospitalization
  • readmission & emergency-return timing proxies

Every removed feature was available only after the outcome was already decided — perfect hindsight, zero clinical value.

Results

The numbers that survive scrutiny.

Held-out Zhang cohort, four horizons, two model families.

0.781

28-day mortality

Random Forest

0.839

3-month mortality

Random Forest

0.710

6-month mortality

XGBoost

0.650

6-month readmission

XGBoost

2026 extension · GCS

Does GCS add anything beyond an ICU stay?

GCS topped the Zhang SHAP ranking, so I retrained inside MIMIC-IV — 42,990 admissions from 18,890 patients, grouped by patient — and tested the components against a real ICU-stay control.

28-day mortality

0.8834 0.8972

+0.0138 AUROC · 95% CI 0.0090–0.0194

6-month mortality

0.8266 0.8352

+0.0086 AUROC · 95% CI 0.0054–0.0120

ICU subgroup, 6-month descriptive

0.8557 0.8766

+0.0208 AUROC · 95% CI 0.0141–0.0274

Paired AUROC differences measured on the same 8,749-admission holdout. Baseline = 34 clinical features; GCS adds eye opening, verbal response and movement.
Outcome Model Features Test AUROC Δ vs baseline (95% CI)
28-day Expanded baseline 34 0.8834 Reference
28-day + GCS components 37 0.8972 +0.0138 (0.0090–0.0194)
28-day + ICU-stay flag 35 0.8915 +0.0080 (0.0037–0.0125)
28-day + ICU flag + GCS 38 0.9000 +0.0166 (0.0119–0.0220)
6-month Expanded baseline 34 0.8266 Reference
6-month + GCS components 37 0.8352 +0.0086 (0.0054–0.0120)
6-month + ICU-stay flag 35 0.8280 +0.0014 (−0.0012–0.0040)
6-month + ICU flag + GCS 38 0.8370 +0.0104 (0.0073–0.0136)
ICU only, 6-month Baseline 34 0.8557 Reference
ICU only, 6-month + GCS components 37 0.8766 +0.0208 (0.0141–0.0274)

Shaded rows: a confidence interval crossing zero (no detectable improvement at this sample size), and the ICU-subgroup result, which uses a single split rather than five-fold cross-validation.

Why the split matters

8,006 of the 18,890 patients have more than one heart-failure admission (one has 59). A row-level split would scatter the same person across train and test and inflate every score. Every number above comes from a grouped 80/20 split on subject_id with zero patient overlap asserted before scoring.

GCS is left missing, never imputed

GCS is charted for 12,083 of 42,990 admissions (28.1%) — the rest were never assessed, so the field is left empty and XGBoost handles the NaN natively. gcs_total is excluded as a feature: for intubated patients it reads 15 regardless of the components. Treating an untested patient as a normal score would fabricate a clinical observation.

The ICU control

The ICU-stay flag comes from the icustays table, independently of whether GCS was charted. At six months the flag alone is a null result (+0.0014, CI crossing zero) while GCS still adds +0.0104 on top of it — so the gain is not simply “was this patient sick enough for intensive care”.

SHAP bar chart of mean absolute SHAP values for the ICU-only MIMIC-IV six-month model
Click to expand
SHAP beeswarm plot for the ICU-only MIMIC-IV six-month model
Click to expand
Scatter of feature rank positions in the Zhang and MIMIC-IV SHAP rankings
Click to expand

Process

Five weeks, four tracks.

Develop, audit, explain, validate — in parallel, not in series.

  1. Step 01

    Develop

    Five candidate models, consistent splits, no peeking.

  2. Step 02

    Audit

    Feature-by-feature leakage review against the data dictionary.

  3. Step 03

    Explain

    SHAP reduction to 143 predictors that carry the signal.

  4. Step 04

    Validate

    Held-out cohort, MIMIC-IV transfer, calibration, DCA, subgroups.

Feb 20–27

Foundations & literature

Feb 28 – Mar 4

Exploratory analysis

Mar 4–8

Preprocessing & features

Mar 9–14

Modelling & comparison

Mar 14–18

Explainability & tool

Mar 18–20

Outreach & dissertation

Explainability

What the final model actually pays attention to.

Mean |SHAP| across the held-out cohort — the cardiorenal axis, in descending order.

GCS (Glasgow Coma Scale)

100%

Moderate–severe CKD

88%

Mitral valve AMS

78%

LV end-diastolic diameter

72%

Liver disease

66%

CHF history

60%

Eye opening (GCS sub-score)

54%

Reduced EF flag

48%

Basophil ratio

42%

Creatine kinase

36%

Why dischargeDay had to go

Day-of-discharge correlates with when the record was written, not with the patient. It inflated every horizon. Removing it cost 0.109 AUROC at six months — and bought a model that means what it says.

A note on GCS

In the Zhang cohort GCS tops the chart partly because it is 15 for 97.2% of patients — its signal is mostly “is it ever low?”. The MIMIC-IV extension answers that properly: leave the score missing instead of imputing a normal one, add an independent ICU-stay flag as a control, and the components still earn a small, consistent increment.

SHAP beeswarm plot for the final 6-month model
Click to expand

External validation

Transportability to MIMIC-IV.

The model was trained on Zhang cohort records and applied to 42,990 MIMIC-IV admissions with no retraining — a transportability check, not a clinical claim.

ROC curves for direct transfer to MIMIC-IV
Click to expand
28-day 0.781 held-out 0.524 direct transfer
6-month 0.710 primary 0.599 conservative, positive
28-day gap
−0.257

Under true domain shift the 28-day AUROC falls from 0.781 to 0.524 — a 0.257 drop that leaves it barely above chance. Reported as a transportability check, not a clinical claim. Retraining on local data is the honest next step, which is exactly what the 2026 GCS extension does.

Evaluation

Calibration, DCA, and subgroup checks.

A model is only as good as its worst-calibrated decile. Brier 0.027, ECE 0.023.

Calibration curves
Click to expand
Decision curve analysis
Click to expand
AUROC heatmap
Click to expand
ROC curves
Click to expand
Kaplan–Meier by NYHA class
Click to expand
Cox forest plot
Click to expand
Subgroup: CKD status
Click to expand
Subgroup: gender
Click to expand
SHAP bar — mean |SHAP|
Click to expand

Glossary

Terms, plainly.

BNP

Range 2.7–5,000 pg/mL in this cohort; danger > 400; mean ≈ 1,280.

LVEF

Healthy 55–70%; reduced < 40%; missing for 68% — imputed with MICE.

NYHA

Class I–IV heart-failure severity; 52% Class III, 31% Class IV here.

AUROC

Chicco & Jurman benchmark ≈ 0.73 for this dataset; final 0.710 at six months.

SHAP

Game-theory attributions: how much each feature moved this prediction.

MICE

Multivariate imputation by chained equations — used on 42 variables.

References

Five sources, no padding.

  1. Zhang et al. (2021). Heart failure clinical records. Scientific Data 8:46. doi:10.1038/s41597-021-00835-9
  2. Goldberger et al. (2000). PhysioBank, PhysioToolkit, and PhysioNet. Circulation 101(23), e215–e220.
  3. Johnson et al. MIMIC-IV. PhysioNet.
  4. Chicco & Jurman (2020). The advantages of the Matthews correlation coefficient. BMC Med Inform Decis Mak 20:16 — baseline AUC ≈ 0.73.
  5. Lundberg et al. (2020). SHAP. Nature Machine Intelligence.