科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-04· bioinformatics

Adding layers of information to scRNA-seq data using pre-trained language models

S. M. Krissmer, J. Menger, J. Rollin, T. M. Vogel, H. Binder, M. Hackenberg

原始摘要(英文原文)· Original abstract
Pre-trained language models promise to enrich single-cell analyses with contextual information from large biomedical text corpora, but it remains unclear how to optimally align this knowledge with quantitative scRNA-seq data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then fine-tune lightweight encoder-only biomedical language models to learn a shared, literature-enriched representation. Controlled evaluations across immune and developmental datasets show that this representation preserves cell identity while adding robust and interpretable contextual layers of functional, disease-associated, and developmental information to single-cell analysis workflows.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Adding layers of information to scRNA-seq data using pre-trained language models — 科研速览 Science Skim