Katja Venko, Eva Žerovnik
Over more than a decade of development, various data-driven computational approaches based on the physicochemical properties of proteins have been developed for estimating the aggregation potential of proteins. The currently available algorithms and models enable a sequence-based design strategy to predict regions prone to transform into liquid-liquid phase-separated (LLPS) condensed or solidified aggregated states; for the latter, one can distinguish amyloid and prion-like aggregates. In this study, we performed a comprehensive in silico experiment; more than 40 models were used to test amino acid sequences of various proteins known to aggregate in neurodegenerative diseases, certain progressive myoclonic epilepsies and mental illnesses. Altogether, 20 proteins were analyzed and their proneness to aggregate or condensate was discussed in view of their possible normal and/or toxic function. The reported large set of computational models enables highly accurate predictions of proteins that are prone to aggregation and/or condensation. Furthermore, such aggregation profiling can be performed for any protein sequence of interest.