Lindeberg P Leite, Teófilo Emidio de Campos, Felipe Campelo
Phylogenetic relationships among organisms define hierarchical structures that can be leveraged for domain adaptation. Traditional approaches in domain adaptation often assume homogeneous source domains or merge heterogeneous sources into a single dataset, which neglects informative inter-domain differences and may lead to negative transfer. In this chapter, a method is presented that explicitly models dependencies among source domains derived from phylogenetic trees, using taxonomy as a proxy for phylogeny and weighting the relative importance of training data across levels for the development of predictive models for linear B-cell epitopes. By capturing these relationships, the approach enhances the adaptability of neural language models and improves generalization across evolutionary branches. Computational results across multiple pathogen taxa indicate consistent performance gains compared to three state-of-the-art baselines, demonstrating the advantages of incorporating phylogenetic information into domain adaptation for epitope prediction.