Orsolya Gereben, Hedvig Tordai, Lana Khamisi, Erda Qorri, Tamás Hegedűs
We developed pLM-SAV, a simple yet effective predictor that leverages protein language models (pLMs). Δ-embeddings, computed as the difference between wild-type and mutant sequence embeddings, are used as input for a convolutional neural network. We trained our model on a well-characterized, labelled set of Eff10k and evaluated it on a non-homologous subset of ClinVar data. This approach performs exceptionally well on the Eff10k test folds and reasonably on ClinVar test sets. On ambiguity-defined subsets, pLM-SAV provides complementary predictions to AlphaMissense and REVEL, although these methods retain higher overall performance on broad ClinVar benchmarks. Our results show that a compact predictor trained on labelled variant-effect data can provide useful predictive performance across multiple benchmarks. Unlike previous methods such as VESPA, pLM-SAV uses no handcrafted features or substitution matrices, relying solely on pLM-derived representations. Δ-embedding-based features may provide useful additional signals for future mutation-effect predictors or integrative models.
MOTIVATION: Predicting whether single amino acid variants (SAVs) in proteins lead to pathogenic outcomes is a critical challenge in molecular biology and precision medicine. Experimental determination of all possible mutation effects is infeasible, and while state-of-the-art tools such as AlphaMissense show promise, their diagnostic performance is insufficient and they are often difficult to run locally.
RESULTS: We developed pLM-SAV, a simple yet effective predictor that leverages protein language models (pLMs). Δ-embeddings, computed as the difference between wild-type and mutant sequence embeddings, are used as input for a convolutional neural network. We trained our model on a well-characterized, labelled set of Eff10k and evaluated it on a non-homologous subset of ClinVar data. This approach performs exceptionally well on the Eff10k test folds and reasonably on ClinVar test sets. On ambiguity-defined subsets, pLM-SAV provides complementary predictions to AlphaMissense and REVEL, although these methods retain higher overall performance on broad ClinVar benchmarks. Our results show that a compact predictor trained on labelled variant-effect data can provide useful predictive performance across multiple benchmarks. Unlike previous methods such as VESPA, pLM-SAV uses no handcrafted features or substitution matrices, relying solely on pLM-derived representations. Δ-embedding-based features may provide useful additional signals for future mutation-effect predictors or integrative models.
AVAILABILITY AND IMPLEMENTATION: The code is available at https://doi.org/10.5281/zenodo.15502498.