Xiaojing Hou, Jinlei Sun, Yunqing Liu, Guoyong Wang
Background: Low sequencing depth causes molecular capture loss and zero inflation in scRNA-seq data, reducing the reliability of downstream analyses. Methods: We propose DepthDiff, a depth-conditional expression enhancement model trained with diffusion-based denoising. Using low-depth expression profiles and sequencing depth ratio as conditions, DepthDiff learns a supervised residual mapping to paired high-depth references. During inference, it requires only a single-step forward prediction. Fixed-UMI downsampling was used to construct benchmarks across three public datasets. Results: DepthDiff outperformed supervised baselines, MAGIC, and unenhanced low-depth data in expression reconstruction and biological signal preservation. Ablation experiments showed that x0 prediction and cosine noise scheduling were important for stability, while full reverse diffusion sampling provided no additional benefit. Cross-dataset transfer and CITE-seq validation supported the generalizability and biological relevance of the recovered signals. Conclusions: DepthDiff is an efficient supervised framework for low-depth scRNA-seq expression enhancement, with gains mainly driven by diffusion-based denoising training rather than generative reverse sampling.