Becky H. Huang, Xun Yan
This Special Issue takes up a timely and pressing question: As generative artificial intelligence (GenAI) tools become increasingly prevalent in educational settings, how are they transforming English language education landscapes? While GenAI offers unprecedented opportunities for language teaching, learning, and assessment, such as personalized instruction, immediate feedback, and access to rich linguistic input, its impact on pedagogical effectiveness and assessment validity remains hotly debated. This collection of studies explored how GenAI is influencing, reshaping, and at times challenging established models of English language instruction and evaluation. We argue that although GenAI holds significant promise for English language education, its value and limitations must be carefully examined through empirical research and critical reflection. We begin by highlighting the Special Issue's key contributions to the field of language education, with a particular focus on the teaching, learning, and assessment of English productive skills. We then turn to the gaps in the existing literature on AI integration in English language education, as revealed by the studies in this issue. These include methodological limitations, assessment-related concerns, and ethical considerations arising from the use of GenAI in current studies. Finally, we reflect on both the promises and potential pitfalls of AI and propose a way forward grounded in a human-centered, research-informed approach aligned with UNESCO's framework (Miao & Holmes, 2023; UNESCO, 2021). Four studies featured in this Special Issue provide rich insights into how GenAI tools are being integrated into TESOL classrooms, both as instructional aids and collaborative learning partners. Jin et al. investigated the use of GenAI as a co-participant in collaborative inquiry tasks aimed at improving students' critical thinking and argumentation (CTA) in writing. Their findings show that when GenAI is used in conjunction with peer collaboration and teacher guidance, it can effectively support the development of CTA. However, the study also revealed concerns among users regarding overreliance on AI and skepticism about the trustworthiness of AI-generated content. These concerns highlight the importance of human scaffolding, such as peer dialogue and critical evaluation, in maximizing the benefits of AI. This aligns with UNESCO's (Miao & Holmes, 2023; UNESCO, 2021) human-centered approach, emphasizing that AI should serve as a complement to, not a substitute for, human interaction in the learning process. Sok and Shin examined how AI-mediated speaking tasks can support L2 learners' oral proficiency. Their study compared two prompt types: aiSUM, which focused on summarization, and aiSPEAK, which offered general speaking practice. They found that learners using aiSUM showed stronger performance in summarization and task focus, suggesting the effectiveness of targeted prompts in guiding AI-supported tasks. Students generally reported positive experiences, noting that ChatGPT's interactive feedback enhanced their engagement and motivation. However, 12% of participants expressed concerns about factual inaccuracies and lack of nuance in the AI's responses, highlighting the need for educators to critically evaluate AI output and provide corrective support. Jiang and Lai explored how GenAI supports digital multimodal composing (DMC) within a process-genre framework. Their findings indicate that students who used GenAI performed better in content structuring, purpose development, and multimodal design than those in non-AI groups. AI tools assisted learners with scriptwriting and storyboarding, enhancing their outline and navigation skills. This study demonstrates how GenAI can scaffold students through complex composing tasks when paired with instructional support. Sim et al. focused on the role of GenAI in teaching L2 pragmatics through role-play. They compared ChatGPT-mediated role-play with traditional peer role-play, finding that AI-based interactions facilitated better short-term learning of speech acts such as requests and refusals. However, peer interactions resulted in greater long-term retention. While AI groups had higher exposure to accurate pragmatic forms, both types of interaction were valued by students. The authors suggest combining AI-mediated and peer-based activities to promote both structured input and sustained learning. Several studies in this issue examine how GenAI influences learner engagement and the quality of learning outcomes. Timpe-Laughlin et al. compared learner interactions in two AI platforms: a Spoken Dialogue System (SDS) and ChatGPT. They found that although both systems fostered oral interactional competence (IC), they did so in different ways. SDS encouraged repair strategies and recipient design – skills important for structured assessments – whereas ChatGPT offered more natural, flexible dialogue suited for authentic conversational practice. The authors advocated for a hybrid approach that combines the affordances of both tools. They concluded that not all GenAI programs are created equal. Educators must be critical consumers, selecting and combining AI tools that best fit their pedagogical goals. Chen et al. explored how English as a foreign language (EFL) university students collaboratively use GenAI, specifically iFLYTEK Spark, for argumentative writing more than a semester, and how this affects critical thinking in argumentation (CTA). They found that students relied on GenAI most during the idea-generation and revision stages, gradually moving from simple fact-finding questions to more focused, critical prompts. Group discussions increasingly involved evaluation, synthesis, and refutation, as students shifted from passively accepting AI output to actively questioning and refining it. In their writing, students' essays showed a wider range of argument elements and a stronger balance between claims and evidence. Overall, students appreciated GenAI for offering diverse perspectives, encouraging them to verify information, and helping them consider alternative viewpoints, while also recognizing that its value depended on thoughtful use, prompt quality, and effective collaboration with peers. Eguchi et al. examined interactional patterns during AI-mediated and human–human role-play tasks. Their results revealed that AI role-plays replicated key features of human interaction, such as turn-taking and structured dialogue. However, AI lacked nonverbal communication elements and demonstrated more frequent overlaps and turn gaps. These findings suggest that while AI can support the assessment of interactional competence, it cannot fully replicate the nuances of face-to-face communication. Kim evaluated the scoring accuracy of ChatGPT-4 on English as a second language (ESL) placement essays under different prompting conditions. The results showed that ChatGPT offered high intrarater reliability and moderate alignment with human scoring. Its performance improved significantly when supplied with detailed rubrics, scoring rationales, and sample essays. While not yet ready to replace human raters, ChatGPT has the potential to serve as a reliable supplementary tool in high-stakes writing assessments. Gjorevski et al. explored ChatGPT's ability to generate criterion-based ratings and feedback. Their study found strong correlations between AI-generated and human scores, as well as detailed, meaningful feedback that closely resembled human teachers' comments. With human oversight, such tools could help enhance formative assessment and classroom-based writing instruction. Jang et al. examined APLUS, an AI-powered feedback system for EFL academic writing. The study found that AI-generated feedback significantly improved students' performance in task fulfillment and organization. However, some students reported lower self-confidence, underscoring the emotional impacts of automated feedback. The authors emphasize the need for models that combine diagnostic feedback with emotional sensitivity, including space for self-assessment and reflective learning to foster learner autonomy. Their findings support a human-centered feedback approach that integrates AI and teacher or peer guidance. Karatay and Xu assessed an SDS-based oral proficiency test modeled after IELTS. Their results showed that the system elicited key features of interactional competence, with strong interrater consistency and reliability. While students generally found the AI competent, some limitations were noted in terms of conversational flow and the absence of nonverbal cues. These findings affirm the value of AI in oral assessment while highlighting the continued need for human involvement. Taken together, the studies in this Special Issue offer an evidence-based understanding of GenAI's role in English language education. Across contexts, GenAI tools consistently demonstrated strong potential for supporting both global and specific productive language skills – including argumentative writing (e.g., Jiang & Lai; Jin et al.), oral communication (e.g., Sok and Shin; Timpe-Laughlin et al.), and critical thinking (e.g., Chen et al.). They also proved effective in generating automatic feedback (Jang et al.) and enhancing assessment validity (Gjorevski et al.; Kim). In assessment, GenAI has also been shown to elicit key features of interactional competence (Karatay and Xu), demonstrate high intrarater reliability and alignment with human scoring (Gjorevski et al.; Karatay and Xu; Kim), and provide detailed, meaningful feedback (Gjorevski et al.; Jang et al.) that improved students' writing performance (Jang et al.). Importantly, these studies also emphasize that not all AI tools are created equal. Timpe-Laughlin et al. and Sok and Shin compared different GenAI systems and found significant variation in outcomes, depending on the design and context of use. Despite its potential, GenAI has notable limitations. To begin with, GenAI tools were less effective in supporting the acquisition of complex linguistic features like delivery and grammar (Sok and Shin). It also lacked the richness of human expression, including the nuances of face-to-face interaction and nonverbal communication (Eguchi et al.; Karatay and Xu; Sim et al.). Additionally, there is a risk of overreliance on these tools, which can negatively impact students' self-confidence and sense of agency (Jang et al.; Jin et al.). While tools like ChatGPT can support language production, the ease of accessing grammatically correct and contextually appropriate text may lead learners to bypass essential processes such as brainstorming, drafting, and revising. As a result, students may engage less critically with language and forgo the cognitive benefits of active learning. This raises the danger of learners becoming passive consumers of AI output rather than active producers of meaning. Overreliance on AI can therefore impede the development of key linguistic and metacognitive skills, weakening learners' long-term language proficiency. GenAI systems are also prone to factual inaccuracies, so users' and educators' critical evaluation and corrective support are essential (Jin et al.; Sok and Shin). These limitations underscore the need for a human-centered approach to GenAI integration, that is, the “human-in-the-lead” rather than “human-in-the-loop” approach. AI should support, not replace, human educators and assessors (e.g., Jang et al.; Timpe-Laughlin et al.). Timpe-Laughlin et al. argue that educators must act as critical consumers, selecting and combining AI tools that align with their pedagogical goals. Similarly, Gjorevski et al. suggest that, with thoughtful oversight, GenAI can enhance formative assessment and classroom-based writing instruction. Jang et al. further support this perspective, reinforcing the central role of human educators in shaping effective AI-supported learning environments. Jiang and Lai, Jin et al., and Sim et al also argue that effective use of GenAI requires deliberate instructional design, ongoing human guidance, and collaboration between humans and AI. Despite this special issue's contribution to an evidence-based understanding of GenAI's role in English language education, it also brings to light several important gaps in the current literature. First, there is a pressing need to investigate the use of GenAI among young learners. All the empirical studies featured in this issue focus exclusively on postsecondary or adult learners, leaving a critical gap in our understanding of GenAI's impact on child language learners. Research to date has also largely examined university students' or adults' experiences with AI-enhanced learning, neglecting the developmental, pedagogical, and assessment considerations distinctive to child learners. Future research should examine the long-term impact of AI-generated feedback on children's language and literacy development. While existing studies have primarily focused on adult learners, little is known about whether AI-supported instructional strategies yield sustained gains in young learners' productive language skills. In particular, there is a need to investigate how teachers can effectively integrate generative AI tools into language and literacy instruction to support children's oral and written language development. The effects of prolonged exposure to AI-mediated instruction on children's linguistic, cognitive, and social development also warrant systematic study, especially given children's distinctive developmental characteristics. Another important direction is the incorporation of AI into the assessment of child learners' productive language skills, not only through automated scoring but also through the design of engaging, gamified, and fully automated assessment environments (Acquah & Katz, 2020; Ma et al., 2025; Yeatman et al., 2024). Such innovations have the potential to transform how educators monitor growth, personalize instruction, and support child learners in educational settings. Second, the growing use of AI-enhanced assessment tools raises concerns about the validity, fairness, and transparency of these practices. While the four studies in this issue that examined AI in assessment found the technology to be generally useful, they also underscore the continued necessity of human oversight. Without rigorous validation and human moderation, AI-based assessments may produce misleading results, complicating efforts to evaluate true learning outcomes. Moreover, the increasing use of GenAI tools by students to generate or refine their responses also poses challenges to language assessment. When it is unclear whether work reflects a student's own language proficiency or the influence of AI, the validity of assessments – especially those targeting productive skills such as writing and speaking – can be significantly compromised. This challenge calls for a critical reexamination of the construct of “language proficiency” in the AI era and how “authentic” language learning should be defined, operationalized and measured (Xi, 2025). Third, several studies raise ethical concerns related to AI use in education, particularly regarding overreliance on GenAI and issues of academic integrity. In a recent study, Ahmad et al. (2023) found that university students' increased use of AI tools contributed to diminished independent thinking, greater passivity, and a worrying level of human “laziness.” While the study does not focus exclusively on GenAI, its findings are highly relevant. Generative tools that automate writing, idea generation, and problem-solving can easily foster overreliance, thereby discouraging the development of learners' critical thinking and communication skills. Similarly, a recent neurocognitive study from MIT Media Lab (Kosmyna et al., 2025) found that students who relied on large language models (LLMs) to write essays showed significantly lower neural engagement compared to those who wrote without digital assistance. The use of LLMs corresponded with decreased brain activity in areas with and suggesting that frequent on GenAI tools may meaningful learning users also reported lower and had of their own writing. these findings to a AI in education, both teachers' and students' to passive consumers of a that concerns about the long-term effects of GenAI on learner and educational integrity. The ease with which GenAI can produce content the risk of and the of questions about and transparency in educational settings. recent from the of these ethical challenges 2025). that had used GenAI tools such as ChatGPT and to while students from using AI. AI-generated and from a and a While the the a about and the importance of regarding AI use by both students and This reflects a among students when this a related study at found that recognizing the benefits of AI, to use it for of social & 2025). Taken together, these findings emphasize the need for and to and for AI use in both teaching, learning, and assessment. As GenAI increasingly in educational these challenges be essential to its and effective the integration of AI in English language education calls for a human-centered approach (Miao & Holmes, 2023; UNESCO, 2021). To this we propose UNESCO's framework of key While may or the framework offers a and that and human we the into four The and of AI use on fairness, and all of which are critical in the context of language education. are in this and that AI serve educational with human in high-stakes of of automated writing evaluation tools should not be the for high-stakes or as such long-term and human oversight. and that AI systems be and for all learners to reinforcing linguistic, or GenAI tools are using from the which related to and In the context of language education, this can in that the language or AI-generated feedback may of English or to specific & 2023; et al., speech tools must also be evaluated for that language learners with or language are not for their diverse & 2023; & 2025; of of The to and the importance of users' the design, and use of AI AI tools such as language learning or speech learner or learners' or writing and must and provide transparency about how the be used et al., 2023; & & 2024). a that learners' speech for feedback should not without users' this that AI-enhanced language teaching and assessment learners' and aligned with ethical and The in this and and and and emphasize the need for and ethical of AI systems in language education. They that human such as AI and and ethical for how AI tools are and used & et al., and how AI systems generate feedback, linguistic features they and they were on et al., The of of (2023) some questions AI users can such as whether or not the AI models is and how well the AI fit educational goals. a writing a student's as grammatically both the and the teacher should be to the or to transparency is the need for and and AI should teacher not replace it – particularly in high-stakes language assessment & of of requires that and AI for system and This AI tools for accuracy different (e.g., of different of models as language use and for users to oversight, and is not only an ethical but also a pedagogical and assessment necessity & et al., The of and calls for and of the of AI systems of of In language education, this students from potential such as or exposure to content AI tools. a an AI tool that responses, must the system cannot be or to access or The of the importance of AI use in language education with and educational of of 2023; et al., AI tools can support learners in or is In in language education requires ongoing evaluation of how AI systems teacher and learner tool that engagement but gradually or peer interaction may not be The two in this emphasize the importance of AI literacy and such as and in AI and for and educational to a understanding of how AI their limitations, and their ethical literature on the understanding that AI literacy is a not only of AI systems but also ethical and critical evaluation et al., 2023; & 2023; & & et al. and et al. (2023) emphasize such as the ability to critically AI's propose et al. evaluation, and while et al. to the and offered a for teacher AI which and and Jiang nuance from a teacher education perspective, underscoring the need for critical engagement with AI like ChatGPT. In this and language teachers not only to use AI tools – such as automated feedback and text – but also to evaluate them teachers should be to when an AI-generated grammar is or when Similarly, learners need to AI literacy to AI metacognitive skills for and of when human feedback is more appropriate than automated important is and which calls for and flexible to AI that learners, and AI of of should for the use of AI in and limitations, and to and academic & 2024). GenAI access to input, which is and to This potential the use of especially when lack for In language education, must reflect the diverse of learners, including learners of different language and learners with learning must also as for and tools, ethical concerns, and that are & 2024). and into language education we can foster not more AI use, but also more and language learning environments of of intelligence a in English language education. GenAI tools offer opportunities to enhance instruction, support personalized learning, and assessment practices. the these tools significant challenges related to fairness, and ethical As this Special Issue has while the potential of GenAI is its in educational must be with than AI as a or it as a to language education, we for a grounded integration of GenAI into language education. This requires sustained empirical ongoing and a to practices. We this special issue not only as an of an evidence-based approach but also as a for further research in this critical engagement with AI is and must work collaboratively to that GenAI is used not for but for educational The of AI in English language education should be that learners, instructional quality, and the of language assessment. we for a human-centered, research-informed approach to GenAI – that both its potential and its limitations. To this AI literacy must become a for all and learners a understanding of how AI and how it can be used we can its as a for and effective in English language education. The authors have that or at the of ChatGPT used during the writing to enhance and The authors carefully and the content and for the of the is a in the of and at The and a at the for Research and research focus on language and literacy development and assessment of students. is a of and and at the of and a at the for and research on the and to language and assessment. not to this as were or during the current