Dianyu Wang, Xiao Li, Zuheng Wang, Chunmeng Wei, Wenhao Lu, Rongbin Zhou, Leifeng Liang, Fubo Wang, Junyi Chen
In this retrospective analysis, we used routinely collected clinical variables to develop a machine learning model for liver cancer diagnosis and examined the variables that contributed to model performance. We studied 3,629 people who were assessed because they were suspected to have liver cancer or other liver diseases. Patients who had pathologically confirmed liver cancer, and controls comprised of healthy subjects, patients with chronic hepatitis B, chronic hepatitis C, or cirrhosis. From 87 clinical parameters, least absolute shrinkage and selection operator regression was applied to identify candidate predictors. Nine machine learning algorithms were then trained and tested through 10-fold cross-validation and external validation in an independent cohort. Performance of the models was measured based on area under the receiver operating characteristic curve, sensitivity, specificity, calibration and decision curve analysis. The variables that contributed most to discrimination included carbohydrate antigen 19-9, alpha-fetoprotein-related indicators, liver function parameters, and inflammatory markers. Among the nine algorithms, Extreme Gradient Boosting achieved the highest discriminative performance, with an AUC of 1.000 in the training cohort and 0.937 in the validation cohort. In the SHapley Additive exPlanations analysis, carbohydrate antigen 19-9 made the largest contribution to the final model. These findings suggest that machine learning algorithms, particularly Extreme Gradient Boosting, may integrate heterogeneous clinical indicators and provide a useful auxiliary tool for distinguishing liver cancer from non-liver cancer conditions in individuals undergoing clinical assessment. Further prospective validation in high-risk surveillance cohorts is required before the model can be applied to early risk prediction or population-level screening.