Tatsuya Hayashi, Shinya Kojima, Norio Hayashi
To quantify how evaluation design affects brain MRI-to-CT synthesis performance, we analyzed 180 MRI/CT pairs from three centers in the SynthRAD2023 Task 1 dataset. A lightweight 2D U-Net was trained with a common validation-based stopping rule under slice-random, patient-random, center-stratified size-matched patient-random, and three center-held-out designs. Comparisons were descriptive. The primary outcome was whole-mask mean absolute error (MAE), supplemented by MAE within reference-CT-defined Hounsfield unit (HU) classes. MAE (mean ± standard deviation) was 94.4 ± 3.0, 100.8 ± 1.3, and 105.2 ± 2.8 HU for the random designs and 187.4 ± 6.0, 160.6 ± 7.1, and 137.4 ± 5.1 HU for centers A, B, and C held out. Slice-random testing yielded lower MAE than patient-random testing. Center-held-out testing yielded higher MAE than the center-stratified comparator, although institutional and acquisition-domain effects could not be separated. Bone errors were high across designs; adipose-range and air errors were elevated in specific center-held-out tests.