科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACM Transactions on Asian and Low-Resource Language Information Processing2026-07-31· Computer science

Large-Scale Word Sense Tagging in Contemporary Japanese: An All-Words Word Sense Disambiguation Approach for 180 Million Words

Kanako Komiya, Soma Asada, Masayuki Asahara

原始摘要(英文原文)· Original abstract
Using a system based on Bidirectional Encoder Representations from Transformers (BERT), the authors automatically performed word sense disambiguation on all content words across multiple Japanese corpora, totaling over 180 million words. A corpus tagged with word senses is searchable based on word senses and can be used to investigate the frequency distribution of each word sense, which is expected to contribute to Japanese language studies and be applied to Japanese language education. However, few Japanese corpora in which word senses are manually tagged are available because annotation of word senses requires specialized knowledge and is costly. The authors developed a word sense disambiguation system using fine-tuned BERT with core data of Balanced Corpus of Contemporary Written Japanese (BCCWJ), manually annotated with word senses. Using this model, the authors assigned word sense tags based on the Word List by Semantic Principles to various Japanese corpora published by the National Institute for Japanese Language and Linguistics. The authors manually evaluated 500 words for each domain in a part of the corpora to which the authors added word senses, and the accuracies were from 68.2%-91.8% according to domains. The authors also introduce a new application for corpus search, Chunagon, which allows searches based on word senses and is now available to the public.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large-Scale Word Sense Tagging in Contemporary Japanese: An All-Words Word Sense Disambiguation Approach for 180 Million Words — 科研速览 Science Skim