科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Data in brief2026-08-01

A parallel English-Akan maternal health dataset to support machine translation and text-to-speech systems.

Isaac Wiafe, Akon Obu Ekpezu, Fiifi Baffoe Payin Winful, Akosua Nyarkoa Wiafe-Akenten, Mark Atta Mensah, Nissi Semanyoh, Justice Kwame Appati, Sumaya Ahmed Salihs, Evans Kwasi, Gifty Odame, Kelvin Nketia-Achiampong, Kwaku Owusu Osei, Frank Ernest Yeboah

原始摘要(英文原文)· Original abstract
Audio and text datasets are essential for developing machine translation, text-to-speech systems, and speech-enabled systems. However, domain-specific datasets for low-resource African languages remain limited, particularly in the healthcare domain. This study addresses this gap by introducing a parallel English-Akan maternal health dataset to support machine translation and text-to-speech systems. The dataset consists of 3487 unique English maternal health phrases (question-and-answer) generated from verified digital health sources and medically reviewed for Ghanaian contextual relevance. These phrases were translated into Akan by four linguistic experts, producing 12,000 Akan transcriptions and 12,000 corresponding Akan audio recordings, totalling 32.313 h. The dataset covers prenatal and postnatal maternal health themes that are aligned with the WHO-recommended domains. The audio recordings were captured in soundproof vocal booths and processed into .wav format. This dataset provides a valuable resource for advancing Akan health communication technologies, maternal health chatbots, and low-resource language AI research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A parallel English-Akan maternal health dataset to support machine translation and text-to-speech systems. — 科研速览 Science Skim