Rahul Kaushik, Suyong Re
INTRODUCTION: The functional diversity of antimicrobial resistance (AMR) proteins necessitates advanced computational frameworks capable of capturing complex macromolecular signatures beyond simple sequence homology. While traditional alignment-based tools identify known resistance determinants, they often overlook the underlying biochemical and structural properties that define novel or divergent resistant macromolecules.
METHODS: This study systematically investigates the capacity of classical sequence-derived descriptors and deep protein language model (pLM) embeddings to resolve these essential functional characteristics at the protein level and contribute to predictive robustness and generalization. Multiple machine learning classifiers were trained using classical features, deep embeddings, and their integrated representations to assess the consistency of AMR-related signal detection.
RESULTS AND DISCUSSION: Across models, both feature types demonstrated strong discriminative capacity while deep embeddings-based detection outperforming the classical features-based detection. Further, their integration did not produce substantial performance gains over deep embeddings alone, however it consistently improved stability and generalization across classifiers and validation settings. These findings indicate that classical descriptors encode resistance-relevant signals that complement contextual representations learned by protein language models. The AUROC values approached 0.85 for classical features alone, 0.97 for deep embeddings alone and 0.94 for integrated feature space across experimental configurations, confirming that alignment-free numerical representations effectively capture AMR-associated functional characteristics. Benchmarking against diverse AMR identification tools further demonstrates that classical and deep representations-based approaches maintain sensitivity toward divergent sequences while preserving high specificity, individually as well as collectively. All datasets, models, and scripts are publicly available at https://github.com/DrRahulKaushik/ProARG.