Minghao Li, Minglei Dong, Hongle Li, Yuanxing Peng, Zhen Li
High-performance computer-aided drug design is a promising field, in which drug-target affinity (DTA) prediction serves as a core step to reduce R&D costs and improve efficiency. However, current DTA methods suffer from insufficient multi-modal representation and lack interpretability. We propose MAIDTA, an interpretable attention-based multi-modal model integrating sequence and graph dual-modal representations of drugs and proteins. An attention-driven binding region module quantifies drug atom-protein residue interaction strength, and a cross-scale feature fusion module integrates multi-scale features. Experiments on Davis and KIBA show MAIDTA outperforms state-of-the-art models with higher accuracy, robustness, and interpretability, offering new insights for drug discovery.