Nikita Murzintcev, Shakhlo Shukurlaevna Yuldasheva
Although Uzbek is the official language of Uzbekistan, it remains one of the low-resource languages. In this paper, we propose a morphology analyser based on the Hunspell library. Such a tool cannot be developed without a detailed description of the Uzbek language, so the main part of the paper is dedicated to studying Uzbek morphology in a way sufficient for building morphology-analysis and spell-checking software. A special emphasis is placed on listing all possible word forms that can be produced with inflectional affixes, accounting for irregular forms, and phonetic assimilation. The orthography of the language is provided in accordance with the latest reforms in the script system and supplemented with transliteration rules and recommendations for text normalization for computer processing. The produced lemmatizer is compatible with a wide range of existing software. The results of this project include a dictionary of Uzbek lemmas annotated with parts of speech tags.