Jai Paris, Oliver Kleinig, Ayushi Agarwal, Weng Onn Chan, Dinesh Selva
Large language models (LLMs) such as ChatGPT, are increasingly used to search the literature and answer focused clinical questions. However, evidence on their reliability for study retrieval remains limited, particularly in Ophthalmology, where conditions are rare, terminology is nuanced, and niche management continually evolves [ 1 ]. Reports of hallucinated references in earlier-generation models underscores the need for contemporary, domain-specific evaluation [ 2 ].