科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ KDD : proceedings. International Conference on Knowledge Discovery & Data Mining2026-08-01

Why Only 3D for 3D? 2D-Guided Supervision for 3D Medical Vision-Language Models.

Qilong Zhao, Ziyuan Qin, Liang Zhao

原始摘要(英文原文)· Original abstract
Training 3D medical vision-language models (VLMs) requires paired volumetric scans and expert annotations such as reports or segmentation masks, which are costly and limited. Meanwhile, large collections of unlabeled 3D scans and strong 2D medical VLMs are increasingly available. We study a simple paradigm: using pretrained 2D models as automatic annotators to supervise 3D VLMs. A 2D teacher generates slice-level descriptions or masks that are aggregated into volume-level pseudo-labels to train a 3D student operating on full volumes at inference. We instantiate this 2D-to-3D supervision framework for report generation and segmentation. In label-scarce regimes, 2D-derived pseudo-labels significantly improve data efficiency. For report generation, pseudo-reports outperform or match ground truth when reports are limited and remain competitive at larger scales. For segmentation, pseudo-masks enable learning from unlabeled volumes and further improve performance when combined with limited expert masks. Overall, strong 2D models can act as scalable annotators that complement scarce 3D labels. This data-centric perspective offers a practical path to scaling 3D medical VLMs by leveraging abundant 2D expertise without changing model architectures.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Why Only 3D for 3D? 2D-Guided Supervision for 3D Medical Vision-Language Models. — 科研速览 Science Skim