科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International journal of molecular sciences2026-09-12

Retrieval-Guided Transfer Learning for Low-Resource Ebola Drug-Target Affinity Prediction.

Mubarakah Alotaibi, Nada Al Taweraqi

原始摘要(英文原文)· Original abstract
Drug-target affinity (DTA) prediction plays an important role in computational drug discovery; however, its application to emerging infectious diseases such as Ebola remains challenging because of the limited availability of experimentally measured affinity data. To address this low-resource setting, we propose a retrieval-guided transfer-learning framework that leverages BindingDB interactions to improve Ebola DTA prediction. The framework uses a two-stage strategy. In Stage I, source interactions are selected using random sampling, compound-similarity retrieval, protein-similarity retrieval, or hybrid compound-protein retrieval at source-data budgets of 50,000 and 300,000 interactions and combined with Ebola training data to learn transferable representations. In Stage II, the pretrained compound and protein encoders are frozen, while the prediction layers are adapted to the Ebola domain. The framework was implemented with DeepDTA and GraphDTA and evaluated across five random seeds using a scaffold-based split, with conventional machine-learning models and single-stage deep learning as baselines. The best overall configuration, GraphDTA with protein-similarity-guided retrieval at the 50,000-interaction budget, achieved RMSE =0.5498±0.1073, R2=0.8809±0.0452, and Pearson =0.9395±0.0237, outperforming the strongest conventional machine-learning model (Extra Trees, RMSE =0.6065±0.0642) and single-stage GraphDTA (RMSE =0.6533±0.0752). Across both source-data budgets and both DTA backbones, all targeted retrieval configurations achieved lower mean RMSE than their corresponding random-retrieval configurations. At the 50,000-interaction budget, matched seed-wise analysis further showed consistent improvements for protein-guided retrieval across all five seeds for both backbones. Retrieval characterization showed stronger target-domain similarity and substantially lower cross-strategy overlap at 50,000 than at 300,000 interactions, while post hoc sequence analysis independently confirmed enrichment of sequence-level relatedness with protein-guided retrieval. Increasing the source-data budget from 50,000 to 300,000 did not uniformly improve predictive performance, indicating that source-data relevance and retrieval selectivity should be considered jointly with source-data quantity. Finally, virtual-screening and approved-drug repurposing case studies across six Ebola virus targets demonstrate the use of the framework for computational prioritization of compound-target hypotheses. Overall, the findings support relevance-guided source-data selection as an effective strategy for transfer learning in low-resource DTA prediction.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Retrieval-Guided Transfer Learning for Low-Resource Ebola Drug-Target Affinity Prediction. — 科研速览 Science Skim