Jinze Huang, Huanyue Liao, Bo Meng, Guangkui Fan, Dong An, Xinhua Dai, Xiang Fang, Yang Zhao
Ensuring robust quality control (QC) remains a major challenge in quantitative proteomics, particularly in detecting and managing outliers. Deep learning offers powerful representational capacity for ultra-high-dimensional data but often suffers from overfitting in small-sample scenarios. To address this, we propose memory-constrained outlier detection (MCOD), a deep anomaly detection framework that directly processes MaxQuant outputs and achieves competitive recall at high precision levels, which suggests a reduced risk of overlooking true outliers. MCOD integrates two innovations: (i) a memory-constrained module (MC module) that mitigates over-representation of samples via prototype-based regularization, and (ii) an adaptive steady-aware regulator that dynamically adjusts the per-sample loss weights in the MC module according to the estimated overfitting risk. Across two simulation settings based on a human cervical cancer cell line (HeLa) proteomics dataset and two real-world cancer proteomics datasets, MCOD consistently outperformed 18 statistical, machine learning, and deep learning baselines, achieving superior area under the receiver operating characteristic curve and area under the precision-recall curve scores. Functional enrichment analyses on the two real-world datasets showed that MCOD performed favorably compared to the three domain-specific models. Furthermore, feature-level visualization provided insights into the rationale behind the model's anomaly assignments. Collectively, MCOD establishes a robust and scalable framework for data QC in quantitative proteomics.