I. Kitsos Kalyvianakis, C. Perez de Amezaga, E. S. Lee, A. Palmer, J. Garner, U. S. Wu, A. Orfanoudaki
Sepsis is a leading cause of hospital mortality, yet timely recognition is hampered by nonspecific presentations and label noise in code-based case definitions. We developed STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality. We refined a pragmatic operational definition using a large language model applied to discharge summaries, with an independent physician-adjudicated cohort as the gold standard. We compared 8-, 24-, and 48-hour observation windows and benchmarked against SOFA, SIRS, and Epic, assessing calibration and discrimination. Among 356,610 encounters, the 8-hour model achieved an AUC of 0.960 in derivation and 0.878 in physician-adjudicated validation, matching or outperforming longer-window models. STRIDE outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters and reaching 78.2% specificity at 80% sensitivity in validation. These findings support accurate sepsis surveillance while limiting unnecessary alerts and requiring minimal prior history.