TianXing Qiao, YongE Feng
Nuclear receptors (NRs) are ligand-dependent transcriptional regulators whose ligand-binding residues are crucial for receptor activation and drug-target analysis. In this study, we developed NRB-FusePS, a protein language model-based multimodal framework for predicting BioLiP-defined ligand-binding residues in NRs. Nuclear receptor ligand-binding information was collected from the BioLiP database, and residue-level samples were generated using 21-residue sliding windows. A protein-level independent test strategy was used to reduce information leakage caused by overlapping windows from the same protein. Sequence composition, amino acid pair, structural, evolutionary, and ESM2 protein language model descriptors were extracted and evaluated using conventional machine learning models. Based on these results, a three-branch deep learning model called NRB-Fuse was constructed to integrate sequence statistical descriptors, structural descriptors, and ESM2 embeddings, and protein-level probability smoothing was introduced to obtain NRB-FusePS. On the independent test set, NRB-FusePS-Full achieved an area under the receiver operating characteristic curve (AUC) of 0.814, a Matthews correlation coefficient (MCC) of 0.406, and a sensitivity of 0.964. Shapley additive explanations (SHAP) analysis showed that ESM2 was the main information source, and the structural case of O00482 supported the predictive utility. The NRBsite web server provides three prediction modes for different levels of input information and is freely available at http://47.95.115.29:8501/ .