科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ npj Digital Medicine2025-12-17· Computer science

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Sana Tonekaboni, Sam Friedman, Xinyi Zhang, Mahnaz Maddah, Caroline Uhler

原始摘要(英文原文)· Original abstract
Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning — 科研速览 Science Skim