Dulanjalee Devage Dona
Artificial intelligence (AI) tools are rapidly used in academic research and industry to generate content and assist with analytical tasks. Because of the recent advances in generative AI, these tools can help with statistical data analysis, including data preprocessing, exploratory analysis, statistical modeling, and interpretation of results. However, concerns remain regarding the reliability and accuracy of AI-generated statistical outcomes and interpretations. This study evaluates the effectiveness and reliability of AI-assisted statistical analysis by comparing the performance of ChatGPT, Claude AI, and Google Gemini using a structured evaluation framework. The analysis was conducted using two datasets from the education and healthcare domains, which are among the most prominent data-driven fields. The study examines how effectively these AI tools select appropriate statistical methods, perform the analysis, and interpret the statistical outputs across multiple dimensions, including statistical accuracy, interpretive quality, domain appropriateness, limitation awareness, and consistency. The study concludes that AI can effectively support statistical analysis but should not replace careful human evaluation. Reliable AI-assisted interpretations still depend on well-structured prompts, critical review, and human oversight. These findings provide practical guidance for researchers, educators, and data practitioners and emphasize the importance of using AI responsibly in both statistics' education and real-world data analysis. Keywords: artificial intelligence, generative AI, large language models (LLMs), statistical analysis, data interpretations