Guillermo Vera-Amaro, José Rafael Rojano-Cáceres
Web accessibility remains a persistent challenge, particularly for visually impaired users who rely on screen readers. This study investigates the potential of large language models (LLMs) to remediate accessibility issues only through structured prompt engineering without accessibility specialization. We evaluate GPT-4o and Gemini 2.0 Flash across 20 variants from two websites using different input formats (HTML, Markdown) and template-guided strategies. Outputs were assessed with automated tools, token efficiency metrics, and manual evaluations with experts and blind users. Results show an average Lighthouse score of 93.25, with WAVE errors reduced by 92.85%. Usability evaluation yielded an average success rate of 95.83% on completed tasks, with accuracy values reaching up to 0.86. GPT-4o demonstrated greater token efficiency, while Gemini produced more visually dynamic outputs. Certain violations persisted, confirming the need for human-in-the-loop validation. Overall, findings suggest that effectively guided LLMs can streamline remediation and foster more inclusive web experiences.