Safkat Mustakim Nirjash, Sharmin Akter, Shakim Ahamed, Emon Hasan
We present PH-AIDE (Public Health AI Decision Engine), a multi-modal, explainable AI framework for five-class intensive care unit (ICU) risk stratification and evidence-based public health policy design. Drawing exclusively on MIMIC-III critical care data, PH-AIDE fuses three complementary modalities from the same patient cohort: (i) structured clinical features (vitals, laboratory values, ICD-9 diagnoses, comorbidity indices) processed by an XGBoost classifier; (ii) hourly vital-sign time-series from the first 48 hours of ICU admission processed by a two-layer bidirectional LSTM (Bi-LSTM) network with additive attention; and (iii) temporally filtered clinical progress and nursing notes (timestamped within the first 48 hours of admission, with discharge summaries explicitly excluded to prevent information leakage) processed by a fine-tuned BioBERT encoder. Module outputs are fused by a two-layer attention-weighted gating network trained in a staged procedure. Evaluated under temporally blocked cross-validation stratified by risk class, PH-AIDE achieves a macro-averaged F1 score of 0.88 ± 0.01, AUC-ROC of 0.93 ± 0.01, and Brier score of 0.081 ± 0.01 across five risk strata, statistically significantly outperforming all single-modality and naïve ensemble baselines (p < 0.01, paired bootstrap, B = 10,000). A comprehensive ablation study confirms the independent contribution of each modality and the superiority of learned attention weighting over mean fusion. On a separate CDC FluSight walk-forward validation, a standalone Bi-LSTM achieves a 1-week-ahead MAE of 0.38% versus 0.51% for ARIMA and 0.41% for the FluSight-ensemble benchmark. An integrated policy simulation layer using constrained linear programming translates probabilistic risk forecasts into SDG-3-aligned resource allocation recommendations, with WHO Global Health Observatory indicators providing cross-national context. Dual-granularity SHAP explainability (global beeswarm + patient-level waterfall) and subgroup-stratified fairness evaluation with multi-class demographic parity difference (macro-DPD ≤ 0.04) are reported. An extended discussion addresses ethical considerations, governance alignment with WHO AI principles, and reproducibility.