Xingxing Li, Jing Jiang, Yugang Dai
Self-supervised learning and multimodal feature fusion approaches have led to significant advances in molecular property prediction. However, most existing multimodal fusion approaches rely on simple feature concatenation, which fails to effectively integrate multi-scale information from different molecular representations. To address this limitation, we propose a novel molecular representation learning framework, termed SMFP, which integrates self-supervised learning and multimodal feature fusion. The proposed framework leverages self-supervised learning to extract feature representations from SMILES sequences and molecular graphs, capturing both local and global molecular information from complementary perspectives, thereby improving the performance of downstream prediction tasks. Specifically, we design a multimodal pretraining framework that promotes feature fusion between SMILES sequences and molecular graphs. This framework incorporates a bidirectional cross-attention module for feature fusion, enabling the framework to better capture the relevance and importance of different molecular modalities and to more effectively integrate multimodal feature information. In addition, a non-overlapping adaptive masking strategy is introduced and applied to tokenized SMILES sequences and graph representations, further encouraging comprehensive feature fusion between the two modalities. Experimental results on eight classification datasets and six regression datasets from MoleculeNet demonstrate that the proposed method achieves competitive performance compared to representative baseline methods, showing advantages on 10 out of the 14 datasets, which validates its effectiveness.