Taichi Endoh, Gerry Amor Camer, Kotetsu Kayama, Daiji Endoh, Hiroki Teraoka
PLM embeddings provide a scalable framework for comparative proteome characterization, complement conventional sequence-based analyses, and prioritize orthologs or protein regions with elevated functional divergence for experimental validation in cross-species pharmacology, toxicology, and systems biology.
BACKGROUND: Comparative proteome analysis can reveal functional conservation and divergence among orthologous proteins, with important implications for pharmacology and toxicology. Protein language models (PLMs) may capture sequence-derived functional relationships beyond what conventional alignment metrics capture.
METHODS: Orthologous proteins from Danio rerio and Danio aesculapii were compared using embeddings generated by the Evolutionary Scale Modeling 2 (ESM-2) protein language model. Reciprocal best-hit inference identified 68,971 high-confidence ortholog pairs, of which 51,086 were available for embedding-based analysis. PLM divergence was quantified using cosine distance and evaluated using length-matched and bitscore-matched random controls, Gene Ontology graph-distance analysis, and localized domain-level comparisons.
RESULTS: Ortholog pairs showed strong global conservation, with a median PLM distance of 0.000487, whereas randomized controls exhibited substantially greater divergence. Increasing Gene Ontology graph distance broadened PLM-distance distributions, and leaf-parent comparisons demonstrated significant functional ordering (Wilcoxon p = 2.44 × 10-4). Local analyses revealed increased divergence in pathophysiologically relevant regions of aryl hydrocarbon receptor (AHR) and potassium channel proteins.
CONCLUSIONS: PLM embeddings provide a scalable framework for comparative proteome characterization, complement conventional sequence-based analyses, and prioritize orthologs or protein regions with elevated functional divergence for experimental validation in cross-species pharmacology, toxicology, and systems biology.