Rui Li, Puyu Zhang, Chenglin Wen, Guoxi Sun
Medical-image annotation is costly and limits the use of fully supervised learning. This study proposes the Self-Supervised Stacked Masked Denoising Autoencoder (S2MDAE) for four-class brain MRI classification. During pre-training, three convolutional encoder-decoder blocks perform layer-wise masked reconstruction under a 75% mask ratio and Gaussian corruption; the pretrained encoder stack is then fine-tuned for classification. On the primary dataset, S2MDAE achieved 87.207% Accuracy, while ablation studies supported the contributions of pre-training, masking, noise injection, and layer-wise reconstruction. Representation analysis further showed increased class separability in deeper blocks after fine-tuning. Without retraining, external evaluation on a duplicate-screened independent dataset achieved 67.254% Accuracy and 66.009% Macro-F1, indicating partial transfer under domain shift. These findings support the task-specific value of the proposed framework while showing that broader clinical generalization requires further validation.