Qilong Zhao, Ziyuan Qin, Liang Zhao
Training 3D medical vision-language models (VLMs) requires paired volumetric scans and expert annotations such as reports or segmentation masks, which are costly and limited. Meanwhile, large collections of unlabeled 3D scans and strong 2D medical VLMs are increasingly available. We study a simple paradigm: using pretrained 2D models as automatic annotators to supervise 3D VLMs. A 2D teacher generates slice-level descriptions or masks that are aggregated into volume-level pseudo-labels to train a 3D student operating on full volumes at inference. We instantiate this 2D-to-3D supervision framework for report generation and segmentation. In label-scarce regimes, 2D-derived pseudo-labels significantly improve data efficiency. For report generation, pseudo-reports outperform or match ground truth when reports are limited and remain competitive at larger scales. For segmentation, pseudo-masks enable learning from unlabeled volumes and further improve performance when combined with limited expert masks. Overall, strong 2D models can act as scalable annotators that complement scarce 3D labels. This data-centric perspective offers a practical path to scaling 3D medical VLMs by leveraging abundant 2D expertise without changing model architectures.