科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Briefings in bioinformatics2026-09-01

ST-ConMa: a multimodal foundation framework for spatial transcriptomics via image-gene contrastive and matching learning.

Taelim Yeom, Dongha Choi, Hyunju Lee

原始摘要(英文原文)· Original abstract
Spatial transcriptomics (ST) enables the spatial mapping of gene expression by combining tissue images with spot-level gene expression profiles, allowing for the analysis of gene distribution and interactions between adjacent regions. Although numerous deep learning models have been developed for ST, most existing approaches are task-specific, and multimodal foundation models with broad knowledge of ST remain scarce. While a few ST multimodal models currently exist, they have been either trained on low-resolution tissue images or developed for a specific platform, thereby limiting their generalizability when applied to diverse ST data. To address this gap, we propose ST-ConMa, a multimodal foundation model pretrained on large-scale spot-level tissue images paired with corresponding gene expression profiles using contrastive and matching learning. While conventional task-specific and foundation models generate modality-specific embeddings for images and genes separately, our model jointly learns image-gene fusion representations using a multimodal encoder equipped with cross-attention and a matching objective. Through extensive evaluations on histopathology image classification benchmarks and domain-specific tasks including gene expression prediction and spatial clustering, we demonstrate that ST-ConMa learns high-quality multimodal representations and consistently outperforms existing methods. These results indicate that ST-ConMa serves as a generalizable backbone for diverse ST approaches.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ST-ConMa: a multimodal foundation framework for spatial transcriptomics via image-gene contrastive and matching learning. — 科研速览 Science Skim