科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Natural language processing.2026-03-01· Computer science

An experimental study on data augmentation techniques for named entity recognition on low-resource domains

Arthur Elwing Torres, Edleno Silva de Moura, Altigran Soares da Silva, Mário A. Nascimento, Filipe Mesquita

原始摘要(英文原文)· Original abstract
Abstract Named Entity Recognition (NER) is a natural language processing task that traditionally relies on supervised learning and annotated data. Acquiring such data is often a challenge, particularly in specialized fields like medical, legal, and financial sectors. Those are commonly referred to as low-resource domains, which comprise long-tail entities, due to the scarcity of available data. To address this, data augmentation techniques are increasingly being employed to generate additional training instances from the original dataset. In this study, we evaluate the effectiveness of two prominent text augmentation techniques, Mention Replacement and Contextual Word Replacement , on two widely used NER models, Bi-LSTM + CRF and BERT. We conduct experiments on three datasets from low-resource domains, and we explore the impact of various combinations of training subset sizes and the number of augmented examples. We not only confirm that data augmentation is particularly beneficial for smaller datasets, but we also demonstrate that there is no universally optimal number of augmented examples, i.e., NER practitioners must experiment with different quantities in order to fine-tune their projects.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

An experimental study on data augmentation techniques for named entity recognition on low-resource domains — 科研速览 Science Skim