Joana Rodrigues da Silva, Carlos Daniel Correa de Oliveira Melo, Eduarda Carvalho Leite, Larisse Santos Mendonça Alves, Camila Pinho E Souza Coelho, Jéssica Luiza Mendonça Albuquerque de Melo, Rafaella Christina do Rego Marques, Julia Maria de Sousa Munduri, André Ferreira Leite, Naile Dame-Teixeira
ML models demonstrated modest ability to estimate short-term within-person changes salivary flow in this dataset. Nevertheless, this exploratory study provides methodological guidance for future salivary prediction research by informing sample-size requirements, candidate variable domains, and the impact of outcome definition on event prevalence and model performance. Larger longitudinal cohorts are needed to develop externally validated models and establish clinically relevant definitions of short-term salivary flow decline.
OBJECTIVE(S): Machine learning (ML) approaches have the potential to capture complex, non-linear relationships and may improve the prediction of salivary flow variation. This study aimed to evaluate the feasibility and internally validate ML models to estimate within-person variation in unstimulated (ΔUSF) and stimulated (ΔSSF) salivary flow over a short-term follow-up period.
DESIGN: Data were derived from the "Brasília 2023 cohort," and included participants with paired salivary flow measurements obtained at baseline and after a mean follow-up of 11 ± 3 months. Participants underwent salivary biochemical profiling, hematological testing, dental examinations, and collection of sociodemographic, anthropometric, dietary, and behavioral data. Machine learning models were developed to predict changes in salivary flow and were internally validated using 10-fold cross-validation.
RESULTS: Seventy-two participants had paired salivary flow measurements. Mean ΔUSF and ΔSSF were -0.07 ± 0.29 mL/min and -0.02 ± 0.40 mL/min, respectively. RF models demonstrated high apparent in-sample prediction; however, internal cross-validation demonstrated limited predictive ability. Similar predictive performance was observed for Random Forest, Elastic Net, and Gradient Boosting, suggesting no clear advantage of one algorithm over another in this dataset. Binary analysis using two clinically anchored threshold showed no meaningful event discrimination.
CONCLUSION: ML models demonstrated modest ability to estimate short-term within-person changes salivary flow in this dataset. Nevertheless, this exploratory study provides methodological guidance for future salivary prediction research by informing sample-size requirements, candidate variable domains, and the impact of outcome definition on event prevalence and model performance. Larger longitudinal cohorts are needed to develop externally validated models and establish clinically relevant definitions of short-term salivary flow decline.