Isaac Samir Wasfy, Jamal Ahmad, Hany Elsegeay, Mohamed F Elebiary, Ahmed Haty, Eman M El-Dydamony, Ahmed Mohamed Soliman, Ahmed Alrefaey, Hesham Abozied, Mohamed Algammal, Hossam A Shouman, Maha M Elzamek, Ahmed Abdel Galil Saleh, Ahmed Fetyan Abdelazim Shafi, Osama Mostafa Abdalla Mohamed
Machine-learning models discriminated between culture-positive and culture-negative sample records in a sample-level internal validation. Culture positivity is not synonymous with symptomatic or clinically significant UTI. The findings support further evaluation as diagnostic risk-estimation tools after urinalysis results are available. They do not establish clinical UTI or the safety or effectiveness of starting, withholding or delaying antibiotics. Patient-grouped external validation with clinical outcomes is required before clinical use.
OBJECTIVES: This study aim to develop, compare and internally validate machine-learning models for predicting urine-culture positivity in patients who had both urinalysis and culture ordered and to explore descriptive probability strata. Post hoc secondary analyses examined age subgroups, the incremental contribution of text-derived features, simpler comparators and calibration.
PATIENTS AND METHODS: Urine culture results are typically unavailable for 24-72 h, creating uncertainty during initial assessment, and machine-learning models may help estimate the probability of culture positivity from routinely collected data. This retrospective study included 2530 urine-sample records originating from three university hospitals. Eligibility was based on paired urinalysis and urine culture records rather than symptom-based diagnostic criteria for urinary tract infection (UTI). Thirteen supervised algorithms were evaluated using a stratified 75:25 sample-record split. This constituted internal validation; records were not grouped by patient or centre because stable cross-centre patient, centre and collection-date identifiers were unavailable.
RESULTS: Several gradient-boosting algorithms showed similar discrimination. CatBoost had the numerically highest test-set AUC of 0.858 (95% CI 0.829-0.892), but its AUC did not differ significantly from gradient boosting or XGBoost. At the reported operating threshold, sensitivity was 0.587 (95% CI 0.513-0.662), specificity 0.930 (0.903-0.951), PPV 0.766 (0.693-0.833) and NPV 0.851 (0.821-0.882). Exploratory probability strata separated records with different observed rates of culture positivity, but their clinical utility and safety were not evaluated.
CONCLUSION: Machine-learning models discriminated between culture-positive and culture-negative sample records in a sample-level internal validation. Culture positivity is not synonymous with symptomatic or clinically significant UTI. The findings support further evaluation as diagnostic risk-estimation tools after urinalysis results are available. They do not establish clinical UTI or the safety or effectiveness of starting, withholding or delaying antibiotics. Patient-grouped external validation with clinical outcomes is required before clinical use.