Safouane Boudakkou, Abdelaaziz El Hibaoui
This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia (arywiki) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikipedia dumps, filtering for script consistency and token density, and applying a Context Injection generation strategy to prevent coreference ambiguity. To ensure high information density, the raw generated output (82,500 pairs) was subjected to a rigorous automated quality assurance pipeline. This pipeline utilized a hierarchical trigger confirmation algorithm to remove repetitive administrative census noise and employed composite embedding based clustering (multilingual-e5-large) for semantic deduplication. The final dataset is formatted as a JSONL file, providing paired instructions and responses alongside their taxonomic categories. This dataset provides a native first culturally grounded resource designed to facilitate the supervised fine tuning of Large Language Models (LLMs) in Maghrebi dialects, circumventing the syntactic limitations of translated English instruction sets.