科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-18· health informatics

Exploring Machine Learning Models to Uncover Pathways in ALS Pathogenesis Using Immunohistochemical Features

J. M. Kuruvilla, O. M. Rifai, J. Longden, J. M. Gregory, M. Vallejo

原始摘要(英文原文)· Original abstract
Hexanucleotide repeat expansions in C9orf72 are the most common genetic cause of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia, yet the predictive information contained within associated inflammatory and proteinopathy markers remains incompletely characterised. Here, we reconstructed five machine-learning-ready datasets from quantitative post-mortem measurements of Iba1, CD68, GFAP, TDP-43 and FUS obtained from 10 C9orf72-ALS cases and 10 controls, building upon the digital pathology study of Rifai et al. Random forest, support vector machine, XGBoost, logistic regression, artificial neural network and ensemble approaches were evaluated, together with preprocessing, hyperparameter optimisation and biomarker-specific feature importance. Under conventional 3-fold cross-validation, model performance varied according to both biomarker and classifier, with no single algorithm consistently achieving the highest sensitivity, specificity and accuracy across all biomarkers. Iba1 showed the strongest random forest performance, reaching 88% sensitivity and 83% specificity. However, when random forest was evaluated using patient-grouped 5-fold cross-validation, performance decreased across all five biomarkers, with reductions of 6.98-14.52% in accuracy, 5.68-18.90% in sensitivity and 2.86-13.04% in specificity relative to conventional cross-validation. Iba1 nevertheless remained the strongest-performing random forest biomarker under patient-grouped evaluation. Feature-importance analysis further showed biomarker-specific predictive profiles, with classification generally distributed across multiple pathological measurements rather than dominated by a single feature. These findings demonstrate that estimates of machine-learning performance in quantitative neuropathology are sensitive to validation design and highlight the importance of patient-level separation when multiple pathological observations originate from the same individuals. Further evaluation using fully nested patient-grouped cross-validation is required to establish comparative model performance on unseen patients.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Exploring Machine Learning Models to Uncover Pathways in ALS Pathogenesis Using Immunohistochemical Features — 科研速览 Science Skim