Kangjie Zheng, Junwei Yang, Siyu Long, Wei Ju, Wei-Ying Ma, Hao Zhou, Ming Zhang
We designed a cross-modal mask-predict pre-training model, CrossMol, to capture semantic associations between modalities with unequal information volumes. The model completes missing 3D structure using higher-level semantic information from another modality, such as SMILES. This allows it to learn cross-modal associations and better understand fine-grained 3D structural information. We also introduce a reweighted distance prediction loss to improve the modelling of short-range structural information. Experiments show that CrossMol achieves large performance gains on multiple downstream molecular tasks, attaining state-of-the-art results.
MOTIVATION: Self-supervised pre-training models for molecular data have demonstrated notable results across many downstream tasks. The inherent multimodal properties of molecules have also motivated efforts to capture information from different modalities. However, current multimodal molecular pre-training models usually treat these modalities as equal and independent, despite differences in their information content. Three-dimensional (3D) molecular structures generally contain finer-grained information than the simplified molecular-input line-entry system (SMILES), which primarily captures higher-level semantic information, such as molecular topology.
RESULTS: We designed a cross-modal mask-predict pre-training model, CrossMol, to capture semantic associations between modalities with unequal information volumes. The model completes missing 3D structure using higher-level semantic information from another modality, such as SMILES. This allows it to learn cross-modal associations and better understand fine-grained 3D structural information. We also introduce a reweighted distance prediction loss to improve the modelling of short-range structural information. Experiments show that CrossMol achieves large performance gains on multiple downstream molecular tasks, attaining state-of-the-art results.
AVAILABILITY AND IMPLEMENTATION: The source code and training data are available at https://github.com/zhengkangjie/crossmol. The code is archived at https://doi.org/10.5281/zenodo.19653204.
SUPPLEMENTARY INFORMATION: Supplementary material contains the task-specific hyperparameters.