Joshuan J Barboza, Oscar Andres Ramirez-Teran, Eduardo Tomás-Alvarado, Carlos A Barba, Julián Santa Cruz-Venegas, Carmen Ayala-Jara, Maicol Cortez-Sandoval, Bryam López Tuesta, Euler Tito Chura, Oscar Alexander Braulio Hernández Rios, Oriana Rivera-Lozada, Cesar Bonilla-Asalde
Background/Objectives: Electrocardiography augmented by artificial intelligence (AI-ECG) has been put forward as an inexpensive, scalable means of identifying a wide range of cardiac disorders, but reported performance differs markedly with the condition targeted, the algorithm, and the reference standard. Our pre-specified primary objective was to estimate the diagnostic accuracy of AI-ECG for heart failure/left ventricular systolic dysfunction (HF/LVSD); accuracy in other cardiac conditions was examined condition by condition, and a cross-condition estimate was computed only as a secondary, descriptive summary. Methods: Following the PRISMA-DTA statement, we systematically reviewed diagnostic test accuracy (DTA) studies indexed in PubMed, Scopus, Web of Science, and Embase from inception to 31 March 2026. Eligible records were primary cohort or case-control diagnostic studies applying an AI algorithm to the ECG against an acceptable reference standard, with a reconstructable 2 × 2 table. Screening, extraction, and QUADAS-2 appraisal were performed independently in duplicate. Sensitivity and specificity were pooled with a bivariate random-effects (Reitsma) model whenever four or more studies shared a target condition, with the corresponding hierarchical summary receiver operating characteristic (HSROC) curve; subgroups with fewer than four studies were synthesised descriptively. Pre-specified sensitivity analyses excluded studies at high risk of bias. Certainty was graded with GRADE for test accuracy. Results: Twenty studies (155,442 ECG-reference pairs) were included and 12 (125,568 ECG-reference pairs) entered the quantitative synthesis. In the pre-specified HF/LVSD analysis (k = 5; n = 101,875), pooled sensitivity was 0.86 (95% confidence interval [CI], 0.80-0.90) and pooled specificity was 0.79 (95% CI, 0.70-0.85), with an HSROC area under the curve (AUC) of 0.89; restricting the analysis to the three studies at low risk of bias gave a sensitivity of 0.83 (95% CI, 0.81-0.85) and a specificity of 0.82 (95% CI, 0.75-0.87). The estimated between-study correlation between logit-sensitivity and logit-specificity was -0.86, indicating a strong implicit threshold effect that the bivariate/HSROC framework accommodates. Descriptive subgroup results indicated high accuracy for atrial fibrillation (median sensitivity 0.96, median specificity 0.94; k = 2) and inconsistent performance across coronary disease and acute coronary syndromes (median sensitivity 0.80, median specificity 0.85; k = 2). The secondary cross-condition estimate (k = 12) was 0.84 (95% CI, 0.77-0.89) for sensitivity and 0.87 (95% CI, 0.79-0.93) for specificity (HSROC-AUC 0.92); because it spans heterogeneous targets it is reported as a descriptive summary only. Overall risk of bias was low in 3/20 studies, unclear in 4/20, and high in 13/20, chiefly in the index test and reference standard domains. The Deeks test indicated funnel asymmetry (p = 0.022), which in DTA reviews commonly reflects heterogeneity in true accuracy rather than publication bias. Conclusions: For HF/LVSD, AI-ECG shows consistent and reproducible accuracy against echocardiography, and this estimate proved robust to the exclusion of studies at high risk of bias. Evidence for other cardiac conditions remains condition specific and less certain. Wider adoption would benefit from prospective external validation against guideline-based reference standards, independent adjudication, and transparent reporting of 2 × 2 data at pre-specified operating points.