科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Biometrics2026-07-01

A tree-based kernel for densities and its applications in clustering DNase-seq profiles.

Yuliang Xu, Kaixuan Luo, Li Ma

原始摘要(英文原文)· Original abstract
Modeling multiple sampling densities within a hierarchical framework enables borrowing of information across samples. These "density random effects" can act as kernels in latent variable models to represent exchangeable subgroups or clusters. A key feature of these kernels is the (functional) covariance they induce, which determines how densities are grouped in mixture models. Our motivating problem is clustering chromatin accessibility profiles from high-throughput DNase-seq experiments to detect transcription factor (TF) binding. TF binding typically produces footprint profiles with spatial patterns, creating long-range dependency across genomic locations. Existing nonparametric hierarchical models impose restrictive covariance assumptions and cannot accommodate such dependencies, often leading to biologically uninformative clusters. We propose a nonparametric density kernel that is flexible enough to capture diverse covariance structures and adapts to various spatial patterns of TF footprints. The kernel specifies dyadic tree splitting probabilities via a multivariate logit-normal model with a sparse precision matrix. Bayesian inference for latent variable models using this kernel is implemented through Gibbs sampling with Pólya-Gamma augmentation. Extensive simulations show that our kernel substantially improves clustering accuracy. We apply the proposed mixture model to DNase-seq data from the Encyclopedia of DNA Elements project, which results in biologically meaningful clusters corresponding to binding events of two common TFs.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A tree-based kernel for densities and its applications in clustering DNase-seq profiles. — 科研速览 Science Skim