Boyang Wu, Yuhang Liu, Yue Cheng, Xiangrong Liu, Quan Zou, Leyi Wei, Lei Xu
We propose scLDM, a generative framework based on latent diffusion models for predicting single-cell perturbation responses. scLDM first compresses high-dimensional gene expression into a compact latent space via a variational autoencoder, followed by a conditional diffusion process to generate post-perturbation states, explicitly guided by pre-perturbation cellular state, cell type, and perturbation type. Systematic evaluations on six datasets across diverse biological settings, spanning pharmacological stimulation, viral infection, helminth infection, genetic perturbations, and multi-species immune response contexts, demonstrate that scLDM achieves superior predictive accuracy compared to state-of-the-art methods. Furthermore, the model exhibits strong interpretability, as the learned perturbation embeddings show high functional alignment with known biological mechanisms. Overall, scLDM provides a robust and biologically consistent strategy for in silico perturbation screening.
MOTIVATION: Accurate prediction of single-cell responses to external stimuli is pivotal for deciphering gene regulatory mechanisms and accelerating data-driven drug discovery. However, effectively capturing the complex, non-linear mapping between intrinsic cell states and external stimuli remains an open problem.
RESULTS: We propose scLDM, a generative framework based on latent diffusion models for predicting single-cell perturbation responses. scLDM first compresses high-dimensional gene expression into a compact latent space via a variational autoencoder, followed by a conditional diffusion process to generate post-perturbation states, explicitly guided by pre-perturbation cellular state, cell type, and perturbation type. Systematic evaluations on six datasets across diverse biological settings, spanning pharmacological stimulation, viral infection, helminth infection, genetic perturbations, and multi-species immune response contexts, demonstrate that scLDM achieves superior predictive accuracy compared to state-of-the-art methods. Furthermore, the model exhibits strong interpretability, as the learned perturbation embeddings show high functional alignment with known biological mechanisms. Overall, scLDM provides a robust and biologically consistent strategy for in silico perturbation screening.
AVAILABILITY: The code is available at https://github.com/samrogers1233/scLDM and archived on Zenodo at https://doi.org/10.5281/zenodo.22658164.
SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.