Kanako Komiya, Soma Asada, Masayuki Asahara
Using a system based on Bidirectional Encoder Representations from Transformers (BERT), the authors automatically performed word sense disambiguation on all content words across multiple Japanese corpora, totaling over 180 million words. A corpus tagged with word senses is searchable based on word senses and can be used to investigate the frequency distribution of each word sense, which is expected to contribute to Japanese language studies and be applied to Japanese language education. However, few Japanese corpora in which word senses are manually tagged are available because annotation of word senses requires specialized knowledge and is costly. The authors developed a word sense disambiguation system using fine-tuned BERT with core data of Balanced Corpus of Contemporary Written Japanese (BCCWJ), manually annotated with word senses. Using this model, the authors assigned word sense tags based on the Word List by Semantic Principles to various Japanese corpora published by the National Institute for Japanese Language and Linguistics. The authors manually evaluated 500 words for each domain in a part of the corpora to which the authors added word senses, and the accuracies were from 68.2%-91.8% according to domains. The authors also introduce a new application for corpus search, Chunagon, which allows searches based on word senses and is now available to the public.