L. S. Pugalenthi, A. P. Collins, N. M. Colomer, S.-K. Marques-Kiderle, C. W. Rodriguez, J. C. H. Chaires, W. A. Motta, J. Filella-Merce, J. J. Li, F. Llanos, N. Zhu, S. Rubio-Guerra, I. Illan-Gala, S. Borrego-Ecija, A. Llado, J. Fortea, A. Lleo, R. Sanchez-Valle, M. L. Henry, M. A. S. Santos, S. M. Grasso
Background The nonfluent/agrammatic (nfv) and logopenic (lv) variants of primary progressive aphasia (PPA) disrupt fluency through distinct underlying neurocognitive mechanisms. Differential diagnosis currently requires hours of cognitive-linguistic testing, with additional barriers for bilingual patients due to a shortage of bilingual service providers and a lack of well-established assessment methods. In English speakers, a promising automated approach for differentiating nfvPPA and lvPPA is to derive speech-timing measures and linguistic features from connected speech as input to machine learning (ML) classification algorithms. To our knowledge, this approach has not been evaluated in the context of bilingualism. Methods Thirty four Catalan Spanish simultaneous bilingual patients (lv = 24, nfv = 10) were asked to describe a picture (Western Aphasia Battery Picnic Scene) in both their dominant and non-dominant language. From the participant's recorded response, we derived four feature sets: speech-timing measures, derived with PRAAT; word-level parameters, derived from corpora; linguistic features, derived with the natural language processing tools SpaCy and CLAN; image-text congruence scores, derived with the vision-language encoder Multilingual-CLIP. Each feature set was fed into classification algorithms for differentiating nfv from lv in participants' non-dominant and dominant samples. Then, we combined each feature set's classifier into an ensemble model. We used the McNemar test to determine the statistical significance of differences in classification performance between responses in the non-dominant and dominant language. Results The best-performing classifier achieved F1 macro scores of 93% (word-level parameters) and 92% (ensemble) in the non-dominant and dominant language, respectively. For all feature sets and ensemble models, classification performance did not significantly differ between the non-dominant and dominant language. Ensemble modeling did not significantly improve classification performance in either language. Conclusions Taking advantage of recent advances in multilingual multimodal machine learning, we accurately differentiate Catalan-Spanish bilingual individuals with nfvPPA and lvPPA using a largely automated, time-efficient (1-2 minutes), and ecologically valid connected-speech-based approach. Future directions include evaluating this approach on larger datasets balanced by PPA subtype, using automated transcriptions of connected speech. Our study represents a step towards addressing current inequities in PPA differential diagnosis for non-English-speaking bilingual speakers. Trial registration Data from the clinical trial NCT05741853 was retrospectively analyzed