Xiaogang An, Meng Jiao, Zhi-Wei Dong, Fei Teng, Li-Na Zhang
3D U-Net underperformed with limited data, whereas RF, with the lowest MAE, offered a practical alternative; the log-linear model showed the narrowest LoA, suggesting greater stability. All models showed substantially reduced predictive performance below 17 mm, where manual planning should take priority. This threshold may serve as an alert for performance degradation, but external validation is required before clinical use.
PURPOSE: To compare three models for predicting mean ovarian dose after ovarian transposition in cervical cancer patients and identify geometric factors affecting prediction reliability.
METHODS: We analyzed 100 ovaries from 51 patients using three models: a log-linear model (dose vs. distance), a random forest (RF) incorporating multiple geometric features, and a 3D U-Net for voxel-wise dose prediction. Models were validated using patient-wise 5-fold cross-validation. Accuracy was assessed by mean absolute error (MAE), median absolute error (MedAE), root mean square error, tolerance rates, and Bland-Altman analysis. A sliding-window residual standard deviation curve determined the distance at which prediction uncertainty increased, with subgroup analyses performed accordingly.
RESULTS: RF achieved the lowest MAE (0.803 Gy) and MedAE (0.365 Gy), followed by the log-linear model (MAE: 0.886 Gy) and 3D U-Net (MAE: 0.890 Gy). The log-linear model showed the narrowest limits of agreement, suggesting greater stability. At ±0.5 Gy tolerance, RF significantly outperformed the log-linear model in our cohort (59% vs. 41%, p < 0.05). A turning point was identified at 17.1 mm. In the normal-geometry subgroup (distance ≥17 mm, n=71), RF showed the highest accuracy and consistency (MAE: 0.390 Gy, 95% LoA: -0.97 to 0.98 Gy). In the difficult-geometry subgroup (distance <17 mm, n = 29), all models exhibited markedly larger errors (MAE >1.8 Gy). For high-risk classification (dose >6 Gy), RF achieved an AUC of 0.956.
CONCLUSION: 3D U-Net underperformed with limited data, whereas RF, with the lowest MAE, offered a practical alternative; the log-linear model showed the narrowest LoA, suggesting greater stability. All models showed substantially reduced predictive performance below 17 mm, where manual planning should take priority. This threshold may serve as an alert for performance degradation, but external validation is required before clinical use.