科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ BMC medical informatics and decision making2026-08-06

The word and the way: strategies for domain-specific BERT pre-training in German medical NLP.

Henry He, Johann Frei, Raphael Schmitt

一句话结论 · In one sentence

ChristBERT establishes a new state-of-the-art for German clinical language modeling. Our findings indicate that the optimal domain adaptation strategy is task-dependent and remains crucial, as adapted models consistently outperformed general-purpose language models in our experiments. To support further research and application in German medical NLP, all developed models are publicly released.

原始摘要(英文原文)· Original abstract
BACKGROUND: Digital healthcare generates vast amounts of clinical texts that hold potential for AI-assisted applications. However, existing German biomedical language models either rely on older architectures or are trained on limited data, which may hinder their performance in real-world settings. METHODS: To explore the impact of domain adaptation strategies in German clinical NLP, we developed a family of domain-specific RoBERTa-based language models, collectively referred to as ChristBERT (Clinical- and Healthcare-Related Issues and Subjects Tuned BERT). To address the lack of large-scale German clinical corpora, we curated a 13.5 GB dataset consisting of scientific publications, clinical texts, and health-related web content. Additionally, we employed data augmentation via translation of English clinical corpora. Three domain adaptation strategies were explored: continued pre-training, pre-training from scratch, and pre-training with domain-specific vocabulary adaptation. RESULTS: The resulting models were evaluated on three medical named entity recognition and two text classification tasks. Our models consistently outperformed four existing general-purpose and medical German models on four out of five tasks. The results demonstrate that the choice of domain adaptation strategy significantly influences downstream task performance. Based on the empirical results, pre-training from scratch is effective for highly specialized clinical texts, whereas continued pre-training is suited for more commonly written medical texts. CONCLUSIONS: ChristBERT establishes a new state-of-the-art for German clinical language modeling. Our findings indicate that the optimal domain adaptation strategy is task-dependent and remains crucial, as adapted models consistently outperformed general-purpose language models in our experiments. To support further research and application in German medical NLP, all developed models are publicly released.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The word and the way: strategies for domain-specific BERT pre-training in German medical NLP. — 科研速览 Science Skim