科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural Processing Letters2025-12-17· Topic model

A Comparative Evaluation of Probabilistic and Transformer-Based Topic Models Across Diverse and Multilingual Text Corpora

Micheal Olalekan Ajinaja, Johnson Tunde Fakoya, Yetunde Esther Ogunwale, John Kolawole Omoniyi, Michael Adeniyi Ibiyomi, Akeem Adekunle Abiona, Damilola Akinola

原始摘要(英文原文)· Original abstract
Topic modeling remains essential for uncovering latent structures in large text corpora, yet performance varies across languages, domains, and document lengths. This study compares five models—Latent Dirichlet Allocation (LDA), Collapsed Gibbs Sampling for LDA, LDA2Vec, Top2Vec, and BERTopic—across three datasets: Hausa news, English short texts (20 Newsgroups), and English long-form corpora (PubMed abstracts and legal case summaries). All models were trained using standardized preprocessing and coherence-based topic optimization under fully reproducible settings. Evaluation combined quantitative metrics ( \({C}_{v}\) coherence and perplexity), expert human assessments ( n = 20), and computational profiling. Results show that BERTopic, powered by the multilingual transformer paraphrase-multilingual-MiniLM-L12-v2 , achieved the highest coherence (0.67) and interpretability across multilingual and domain-specific corpora but required greater computational resources. LDA with Gibbs Sampling offered a more efficient alternative with competitive coherence on smaller datasets. Overall, the findings reveal a clear trade-off between semantic depth and computational scalability, providing practical guidance for multilingual and domain-specific topic modeling.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Comparative Evaluation of Probabilistic and Transformer-Based Topic Models Across Diverse and Multilingual Text Corpora — 科研速览 Science Skim