科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Bioinformatics2026-06-17· Bottleneck

Fitness translocation: improving variant effect prediction with biologically-grounded data augmentation

Adrien Mialland, Shuzo Fukunaga, Riku Katsuki, Yali Dong, Hideki Yamaguchi, Yutaka Saitō

原始摘要(英文原文)· Original abstract
MOTIVATION: Data scarcity limits the characterization of protein fitness landscapes and the development of accurate variant effect prediction models. To address this challenge, we introduce fitness translocation, a data augmentation strategy that generates synthetic variants for a target protein by leveraging variant fitness data previously measured in homologous proteins. Using embeddings from protein language models, the method computes the difference between each homolog variant and its wild type and applies these offsets to the target wild-type embedding to create synthetic variants in embedding space. RESULTS: We illustrate the utility of fitness translocation in the context of variant effect prediction on three protein families: IGPS, GFP, and SARS-CoV-2 spike proteins, across different models and training data sizes. Fitness translocation consistently improves predictive performance, particularly under limited training data, and is effective even when augmenting with remote homologs sharing as little as 35% sequence identity. These results illustrate how biologically grounded data augmentation can expand and diversify protein fitness landscapes, supporting more data-efficient protein engineering. AVAILABILITY AND IMPLEMENTATION: The code and datasets are available at https://github.com/adrienmialland/ProtFitTrans.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Fitness translocation: improving variant effect prediction with biologically-grounded data augmentation — 科研速览 Science Skim