Eungi Kim, Frankline Kipchumba, Sein Min
This study evaluates digital object identifier (DOI) hallucination in large language model (LLM)-generated scholarly citations, with a focus on systematic geographic disparities. To conduct this study, we systematically evaluated four LLMs (GPT-4o-mini, Claude-3-haiku, Gemini-2.0-flash-lite, and DeepSeek V3) using standardized information behavior prompts across ten countries with diverse income levels. The models generated 3451 citations, which we validated using the CrossRef API. The results showed that DOI hallucination follows systematic patterns influenced by model choice, geographic context, and publication recency. Hallucination rates exceeded 80% in lower-income countries and increased sharply for publications from the 2020s across all regions. Fabricated citations—citations that appear structurally complete but contain invalid DOIs—were especially prevalent in countries such as India and Bangladesh. Model-specific factors showed the strongest association with hallucination, followed by income level and publication period. These findings raise concerns about the epistemic reliability of LLM-generated scholarly references and underscore the need for region-aware training, real-time DOI validation, and robust verification protocols in academic contexts.