Victor Miguel Hernández-Maldonado
This retrospective study analyzes 219,023 records from Mexico’s open respiratory surveillance data to identify factors associated with hospitalization and evaluate hospitalization risk stratification using interpretable machine learning. After data quality review and preprocessing, 218,557 complete records were analyzed. Logistic regression was used to estimate interpretable odds ratios, threshold tuning was applied to improve recall, and Random Forest was used as a comparative predictive model. Pneumonia was the strongest factor associated with hospitalization, followed by COPD, diabetes, and hypertension. The baseline logistic regression model achieved 76.7% accuracy and 55.3% recall; threshold tuning increased recall to 71.1%, reducing false negatives. Random Forest achieved the best overall predictive balance, with 79.5% accuracy, 71.9% recall, and a 76.9% F1-score. These findings support the use of open surveillance data for interpretable public health analytics and future respiratory hospitalization monitoring.