Pathamakorn Netayawijit, Wirapong Chansanam, Kanda Sorn-in
Accurate used-vehicle price prediction is essential for consumers, dealers, and financial institutions, as pricing dynamics involve complex and non-linear relationships influenced by vehicle condition, depreciation patterns, and heterogeneous market factors. While ensemble learning models have demonstrated strong predictive capabilities, existing studies rarely compare them systematically with linear regression under multiple data partitioning strategies. This study proposes an Optimized Data Partitioning Framework (ODPF) to evaluate model performance and stability across four train–test split ratios (50–50, 60–40, 70–30, and 80–20) using a leakage-free preprocessing pipeline. The framework incorporates a variance-based stability index to quantify the effect of sampling variability, a methodological dimension largely absent from prior vehicle-pricing research. Six algorithms (Linear Regression, Decision Tree, Support Vector Regression (SVR), Random Forest, XGBoost, and LightGBM) were evaluated under consistent preprocessing and experimental conditions. The results indicate that ensemble methods outperform Linear Regression across all evaluation metrics, and Random Forest demonstrated the strongest performance, with a Root Mean Square Error (RMSE) and a coefficient of determination (R²) of 274.26 and 0.9995, respectively. XGBoost and LightGBM also exhibited high accuracy (R² > 0.998), whereas SVR showed limited generalization (RMSE = 1232.73; R² = 0.0800) for the sales analytics dataset. Stability analysis across repeated sampling identifies the 80–20 split as the most reliable configuration, exhibiting lower performance variance and stronger generalization consistency. Overall, the findings indicate that algorithm selection has a greater influence on predictive accuracy than the partition ratio alone, providing practical guidance for developing robust pricing models in the automotive domain.