Maksud Sharipov, Lola Kurbanova, Ruzikajon Kurbonova, Elmurod Kuriyozov
In this article, we present a morphologically annotated lexical dataset designed to support the translation of texts between the Khorezm dialect of Uzbek and standard Uzbek. The database consists of 1445 Khorezm dialect-standard Uzbek word pairs. For each entry, the dialect form, its corresponding standard form, and part of speech are provided, along with a morphological structure segmented into affixes according to such grammatical categories as number (singular/plural), possessive, and person/number properties. During the tagging process, the grammatical system of Uzbek and the specific inflectional properties of the Khorezm dialect were taken into account, resulting in a clear and machine-processable layer of correspondences between the dialect and the standard language. The resulting lexical resource is intended to serve as additional input data for AI-based translation and normalization models, as well as in applications for morphological analysis, spelling and grammar checking, and educational tools for teaching the Khorezm dialect. This dataset constitutes the first systematic lexical-morphological resource for corpus-based research on the Khorezm dialect and lays the groundwork for studies on machine translation and automatic alignment between dialects and the standard language in low-resource Turkic varieties.