科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Data in brief2026-08-01

Dasinya: A dataset and preprocessing pipeline for the Badini Kurdish dialect.

Vaman A Saeed, Karwan Jacksi

原始摘要(英文原文)· Original abstract
Despite notable advances in NLP for low-resource languages, dialects written in non-Latin scripts - especially those without standardised digital resources - continue to receive insufficient attention. The Badini Kurdish dialect, written in Perso-Arabic script and spoken across the Kurdistan Region of Iraq, lacks any dedicated NLP preprocessing pipeline in the literature. This paper presents Dasinya, the first openly released Badini Kurdish text corpus, together with the Badini Dialect Processing Toolkit (BDPT) - the first NLP preprocessing pipeline built specifically for this dialect. Dasinya contains 107 files, 87,545 sentences, and roughly 1.28 million words, spanning two source types and nine sub-genres: six book sub-genres (literary, historical, cultural, scientific, psychological, and educational) and three non-book sub-genres (news, academic/linguistic, and broadcast). These files were drawn from two repositories: the Badirxanian Public Library in Duhok (98 books) and the Semantic Web Lab at the University of Zakho (4 books and 5 non-book files). BDPT is structured as a six-stage pipeline that handles validation, Unicode normalisation - enforcing Badini-specific character mappings and preserving the Zero-Width Non-Joiner as a morpheme boundary marker - structural cleaning, deep cleaning, text preprocessing, and sentence segmentation. Orthographic problems particular to Badini Kurdish are addressed at each stage; none of these are handled by existing Arabic, Persian, or Sorani tools. At Stage 6, sentence segmentation reaches a specialist-validated Acc% of 91.4% across all 107 files, measured under a two-metric framework (OK% and Acc%) with genre-sensitive thresholds that reflect genuine linguistic differences between books, news, broadcast, and academic texts. Both the Dasinya corpus and the BDPT pipeline are openly available at https://doi.org/10.5281/zenodo.20729060 and are intended to underpin future Badini Kurdish NLP work, including tokenisation, part-of-speech tagging, stemming, and stopword removal.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Dasinya: A dataset and preprocessing pipeline for the Badini Kurdish dialect. — 科研速览 Science Skim