Hui Liu, Chuangqi Li, Ziyu Liu, Shuxiao Yang, Hao Wang, Pengrong Yan
This study presents a multimodal pedagogical intelligent agent integrating a fine-tuned LLAMA model for structured content generation, real-time voice cloning for personalized narration, and an enhanced Wav2Lip pipeline for precise audiovisual synchronization. The agent supports interactive question answering, automatic slide generation, and synchronized video instruction with configurable voice and pacing. Its educational impact was evaluated using a three-factor mixed-design experiment with 300 students from Grades 5, 8, and 11 across urban and rural contexts. After a three-week mathematics unit, results showed significant learning gains for primary and junior students, reduced score dispersion, and modest narrowing of the urban–rural achievement gap, alongside higher usability and immersion ratings than traditional search-and-read methods. In contrast, Grade 11 students showed performance declines, attributed to limited content depth, cognitive overload, and reduced learner trust. These findings highlight both the promise and limitations of AI pedagogical agents across school stages.