Hoda Helmy, Abdullah Hosseini, Ahmed Ibrahim, Asfand Baig-Mirza, Ahmed‐Ramadan Sadek, Ahmed Serag
Automated radiology report generation remains a key challenge in clinical AI, requiring alignment between visual, anatomical, and linguistic representations. This study introduces SPINE – a Segmentation-guided Processing and Integration framework for spinal MRI report generation – designed to integrate multimodal (T1-, T2-weighted, and segmentation) inputs for anatomically informed narrative synthesis. Using axial and sagittal MRI datasets comprising 515 and 190 patients, respectively, SPINE was evaluated across multiple multimodal and segmentation-aware configurations. Results show that incorporating segmentation maps substantially improves report accuracy, coherence, and clinical relevance, with the T1 + T2 + Seg model achieving the most balanced performance across all metrics. Models trained on LLM-standardized reports outperformed those trained on human-written text, demonstrating the importance of linguistic consistency. Comparative analysis also revealed that GPT-4o generated more fluent and clinically detailed reports than Grok-3. These findings highlight the potential of combining multimodal imaging with language-model-driven structuring to enhance both the factual accuracy and clinical reliability of automated radiology reporting, paving the way for more interpretable and trustworthy AI-assisted diagnostics.