Mohammad Haekal, Siti Nurul Khotimah, Galih Restu Fardian Suwandi, Freddy Haryanto
Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but its reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift. This study developed an interpretable RR-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error and agreement analyses. The model showed high discrimination in both development and external evaluation, with external ROC-AUC of 0.9868 and PR-AUC of 0.9872. However, external calibration deteriorated despite preserved ranking performance, with Brier score of 0.0654, ECE(15) of 0.1219, calibration intercept of 2.2402, and calibration slope of 1.4879. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with mean absolute error of 0.1473 versus 0.0836, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased from 0.1464 without recalibration to 0.0692 after Platt recalibration and 0.0604 after isotonic recalibration, while median ECE(15) decreased from 0.1224 to 0.0394 and 0.0287, respectively. These findings indicate that RR-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.