科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ La Clinica terapeutica2026-01-01

Machine learning-based classification model to distinguish tumor and tumor-free samples using synthetic plasma proteomic dataset.

Sara Feizyab, Gabriele Bonetti, Alessandro Macchia, Kevin Donato, Immacolata De Luca, Jan Miertuš, Ahmad Jainul Abidin, Pietro Chiurazzi, Ornela Gordani, Xhilda Dhamo, Eglantina Kalluci, Dominika Vešelényiová, Stanislav Miertuš, Iveta Dirgová Luptáková, Jiří Pospíchal, Daniele Generali, Matteo Bertelli

一句话结论 · In one sentence

This in silico proof-of-concept suggests that integrating mass spectrometry-inspired peptide detection with machine learning could support future plasma-based screening and recurrence monitoring. Challenges include plasma proteome complexity and the low abundance of tumor-derived peptides; real-world validation is essential for clinical transition.

原始摘要(英文原文)· Original abstract
BACKGROUND: Tumor-Specific Peptides (TSPs) and Tumor-associated Overexpressed Proteins (TOPs) are promising biomarkers for cancer diagnosis and monitoring. TSPs arise from somatic mutations unique to tumor cells. Their detection in plasma through mass spectrometry-based proteomics provides a non-invasive alternative to traditional biopsy-based diagnostics. We developed a machine learning workflow to classify tumor versus tumor-free samples using synthetic proteomic features. METHODS: A synthetic (in silico) dataset was generated, parameterized using curated proteogenomic resources (CAPD, CPTAC, dbPepNeo2.0) and a literature-based search. For model development, 52 TSP detection indicators and 5 TOP concentrations (AFP, CEA, PSA, VEGF, ADFP) were simulated for 400 individuals, equally distributed between tumor-affected and healthy. No real patient-level data or biological samples were used. Multiple supervised classifiers were trained and validated using 5-fold stratified cross-validation (random seed = 42). RESULTS: In cross-validation, Support Vector Machine (SVM) achieved the highest mean recall (0.88 ± 0.10) and AUC (0.962 ± 0.029) and was selected as the Stage-2 classifier. On the holdout test set (n=80), the two-stage TSP gate + Stage-2 ML strategy achieved accuracy 0.8875, recall 0.90, and a composite-score AUC of 0.924 (score=1.0 if any TSP is detected; otherwise, the Stage-2 predicted probability). CONCLUSION: This in silico proof-of-concept suggests that integrating mass spectrometry-inspired peptide detection with machine learning could support future plasma-based screening and recurrence monitoring. Challenges include plasma proteome complexity and the low abundance of tumor-derived peptides; real-world validation is essential for clinical transition.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Machine learning-based classification model to distinguish tumor and tumor-free samples using synthetic plasma proteomic dataset. — 科研速览 Science Skim