Yuan Tao, Zhexuan Yu, Tianxue Chen, Wei Jin, Weifeng Jin, Li Yu
Phytomedicine extraction-process optimization often relies on regression or machine learning models, but model choice is commonly judged by goodness of fit alone. Here, we assembled 1,148 structurally eligible extraction datasets from systematic literature screening and evaluated seven candidate models within a unified benchmark. The models included full quadratic regression, two subset-selected parsimonious regression models, quadratic ridge regression, support vector regression, partial least-squares regression, and Gaussian process regression. Performance was assessed across fitting, leave-one-out cross-validation, diagnostic behavior, and stability of model-predicted optima. Parsimonious regression models showed the most balanced profile, whereas some machine learning models reduced training errors without consistently improving cross-validated prediction or optimization stability. These results indicate that model selection can itself introduce optimization uncertainty and should be evaluated as a decision-oriented component of extraction-process modeling rather than by goodness of fit alone.