Minkyu Kim, Tae-An Yoo, Ji‐Bum Chung
AI is spreading rapidly, creating significant social and economic value while also raising concerns about its high energy use and environmental sustainability. While prior studies have predominantly focused on the energy-intensive nature of the training phase, the cumulative environmental footprint generated during large-scale service operations, particularly in the inference phase, has received comparatively less attention and remains difficult to compare across studies due to inconsistent boundaries and disclosure practices. To bridge this gap, this study conducts a scoping review of methodologies and evidence on AI carbon footprint assessment. We analyze the classification and standardization status of existing AI carbon measurement tools and methodologies, and comparatively examine the environmental impacts arising from both training and inference stages, and synthesize inference evaluation into a benchmark-oriented framework. In addition, we identify how multidimensional factors such as model size, prompt complexity, serving environments, and system boundary definitions shape the resulting carbon footprint. Our review reveals critical limitations in current AI carbon accounting practices, including methodological inconsistencies, technology-specific biases, and insufficient attention to end-to-end system perspectives. Building on these insights, we propose future research and governance directions: (1) establishing standardized and transparent universal measurement protocols with clear boundary, functional-unit, and disclosure requirements, (2) designing dynamic evaluation frameworks that incorporate user behavior and service-level operating conditions, (3) developing life-cycle monitoring systems that encompass embodied emissions, and (4) advancing multidimensional sustainability assessment framework that balance model performance with environmental efficiency. This paper provides a foundation for interdisciplinary dialogue aimed at building a sustainable AI ecosystem and offers a baseline guideline for researchers and policymakers seeking to evaluate AI environmental impacts across technical, social, and operational dimensions.