Jaehee Chun, McKell Woodland, Austin Castelo, Caleb O'Connor, Mais Altaie, Aashish Gupta, Shanli Ding, Allyson Nguyen, Eugene J Koay, Kristy K Brock
Incorporating report-derived clinical context improves the longitudinal robustness of personalized liver tumor segmentation, particularly in patients with dynamic tumor evolution. These findings suggest that clinically grounded language representations can stabilize personalization across extended follow-up while enabling fully automated and scalable deployment.
PURPOSE: Personalized segmentation has been shown to improve performance over generalized models in follow-up imaging; however, its robustness across large temporal gaps remains unclear. We propose a fully automated, context-informed personalization framework that integrates radiology report-derived clinical context with imaging features to improve longitudinal liver tumor segmentation.
METHODS: This retrospective study included 171 patients for generalized model training and 41 longitudinal patients with follow-up CT scans acquired 3-15 months after baseline for personalization and evaluation. We developed a U-Net-based vision-language segmentation framework that automatically summarizes free-text radiology reports, encodes them into language embeddings, and integrates them into intermediate image features to provide clinical context. A generalized model was trained using paired CT-report data. For each longitudinal patient, personalization was performed at the 3-month follow-up using either image-only input (vision-only personalization) or paired image-report input (context-informed personalization). Models were then evaluated on subsequent 6-15-month follow-up scans without further adaptation. Patients were stratified into Stable and Dynamic cohorts based on relative GTV volume variability. Performance was assessed using the Dice similarity coefficient (DSC) and LogPenalty Score (LPS), with additional ablation studies examining language models and report components.
RESULTS: In the Dynamic cohort, vision-only personalization exhibited diminishing gains at later follow-up, whereas context-informed personalization maintained statistically significant improvements in ΔLPS at 9-15 months (p < 0.05). In the Stable cohort, both personalization strategies showed sustained performance with comparable segmentation accuracy, indicating that incorporating report-derived clinical context did not degrade performance in cases with limited anatomical variability. Ablation studies demonstrated that domain-specific encoder-based language models (BioClinicalBERT) outperformed larger generative models and that effective context integration did not require complex architectural modifications.
CONCLUSION: Incorporating report-derived clinical context improves the longitudinal robustness of personalized liver tumor segmentation, particularly in patients with dynamic tumor evolution. These findings suggest that clinically grounded language representations can stabilize personalization across extended follow-up while enabling fully automated and scalable deployment.