Marcos Martínez Galindo, Marco Luca Sbodio, Mykhaylo Zayats, Rodrigo Ordonez-Hurtado, Raúl Fernández-Díaz, Vanessa Lopez Garcia, Hoang Thanh Lam
The number of unimodal molecule representation constantly increases, and researchers investigate how to combine them. Intuitively, multimodal representations may provide complementary information and combining them promises better performance. In this work, we systematically explore how combining multiple molecular modalities affects the performance of downstream prediction tasks, providing a baseline for informed decision making. Our study covers 7 molecular modalities and combines them using intermediate and late fusion, and 2 neural network architectures (with or without using knowledge graphs). We conduct experiments with 3 benchmarks for drug-target binding affinity, and 22 molecule property prediction. In total, we train and evaluate over 1400 models. In summary, our results show that combining multiple modalities improve the performance provided that effective fusion strategies are chosen. Knowledge-enhanced representation learning further boosts model performance. Notably, we find that even the use of simple late-fusion approaches establishes state-of-the-art results for some tasks. Here, the authors examine whether combining multiple molecular data types improves drug discovery models, finding multimodal approaches boost predictions with effective fusion, and even simple late-fusion methods can reach state-of-the-art performance.