科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ Mendeley Data2026-08-05· Bengali

A Balanced Six-Region Dataset for Bangla Dialect Identification

Taznim Sadab, Md. Tanzim Shahriar, Adib Hassan, Khandoker Nosiba Arifin, Tasmiah Rahman

原始摘要(英文原文)· Original abstract
The dataset is provided as a single tabular file (.xlsx/.csv) containing four columns: Standard English, Standard Bangla, Regional Dialect, and Label. The Standard English and Standard Bangla columns contain the original elicitation prompt sentence used to collect the corresponding regional-dialect translation. The Regional Dialect column contains the sentence as rendered by a native speaker in one of the six target regional dialects. The Label column contains the corresponding region name (Rangpur, Bogura, Sylhet, Chittagong, Pabna, or Cumilla), indicating the dialect class of that sentence. The dataset contains exactly 3,600 sentences in total, with 600 sentences per region. The dataset is provided in randomly shuffled order to avoid any ordering bias. All text is UTF-8 encoded to preserve the Bangla script and dialect-specific spelling conventions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Balanced Six-Region Dataset for Bangla Dialect Identification — 科研速览 Science Skim