科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Canadian Journal of Information and Library Science2026-06-18· Labelling

Generative AI for automatic topic labelling

Diego Kozlowski, Carolina Pradier, Pierre Benz

原始摘要(英文原文)· Original abstract
Topic modelling has become a prominent tool for the study of scientific fields, as they allow for a large-scale interpretation of research trends. Nevertheless, the output of these models is structured as a list of keywords, which requires a manual interpretation for the labelling. This paper proposes to assess the reliability of three LLMs, namely flan, GPT-4o, and GPT-4 mini for topic labelling. Drawing on previous research leveraging BERTopic, we generate topics from a dataset of all the scientific articles (n=34,797) authored by all biology professors in Switzerland between 2008 and 2020, as recorded in the Web of Science database. We assess the output of the three models both quantitatively and qualitatively and measure the effect of the temperature parameter in GPT models and find that, first, both GPT models are capable of correctly and precisely labelling topics from the models' output keywords at the default temperature. Second, 3-word labels are preferable to grasp the complexity of research topics.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Generative AI for automatic topic labelling — 科研速览 Science Skim