科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Data in brief2026-08-01

AryWiki-Instruct: A high-fidelity instruction tuning dataset for Moroccan Arabic (Darija).

Safouane Boudakkou, Abdelaaziz El Hibaoui

原始摘要(英文原文)· Original abstract
This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia (arywiki) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikipedia dumps, filtering for script consistency and token density, and applying a Context Injection generation strategy to prevent coreference ambiguity. To ensure high information density, the raw generated output (82,500 pairs) was subjected to a rigorous automated quality assurance pipeline. This pipeline utilized a hierarchical trigger confirmation algorithm to remove repetitive administrative census noise and employed composite embedding based clustering (multilingual-e5-large) for semantic deduplication. The final dataset is formatted as a JSONL file, providing paired instructions and responses alongside their taxonomic categories. This dataset provides a native first culturally grounded resource designed to facilitate the supervised fine tuning of Large Language Models (LLMs) in Maghrebi dialects, circumventing the syntactic limitations of translated English instruction sets.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AryWiki-Instruct: A high-fidelity instruction tuning dataset for Moroccan Arabic (Darija). — 科研速览 Science Skim