Taznim Sadab, Md. Tanzim Shahriar, Adib Hassan, Khandoker Nosiba Arifin, Tasmiah Rahman
The dataset is provided as a single tabular file (.xlsx/.csv) containing four columns: Standard English, Standard Bangla, Regional Dialect, and Label. The Standard English and Standard Bangla columns contain the original elicitation prompt sentence used to collect the corresponding regional-dialect translation. The Regional Dialect column contains the sentence as rendered by a native speaker in one of the six target regional dialects. The Label column contains the corresponding region name (Rangpur, Bogura, Sylhet, Chittagong, Pabna, or Cumilla), indicating the dialect class of that sentence. The dataset contains exactly 3,600 sentences in total, with 600 sentences per region. The dataset is provided in randomly shuffled order to avoid any ordering bias. All text is UTF-8 encoded to preserve the Bangla script and dialect-specific spelling conventions.