Hlib Aleksandrenko, Martin Komenda
National health information portals can integrate generative AI with RAG architecture to deliver verified, source-grounded health information, though trade-offs in user-oriented dimensions warrant consideration in future implementations. Still, this expert evaluation assessed response characteristics rather than direct effects on health literacy, awareness, or health outcomes.
BACKGROUND: Generative AI (GenAI) shows promise for health literacy and awareness, with large language models (LLMs) increasingly used for health information-seeking despite concerns about misinformation. National health information portals provide authoritative, evidence-based content but lack conversational interfaces. Retrieval-augmented generation (RAG) offers a potential solution by grounding responses in verified health information portal content, yet evidence on such integration remains limited.
OBJECTIVE: This study aimed to design a national health information portal-integrated generative AI chatbot to support access to trustworthy health information and to evaluate its response characteristics relative to general-purpose LLMs through a multidimensional expert assessment.
METHODS: A RAG-enhanced chatbot (NZIP AI assistant) was designed for the Czech Republic's National Health Information Portal and evaluated by eleven experts who compared it with ChatGPT, Microsoft Copilot, and Google Gemini across 67 health-related questions. Responses were rated on trustworthiness, clarity, user-friendliness, and relevance using a 5-point Likert scale (1 = best, 5 = worst). Content analysis examined response characteristics.
RESULTS: In total, 267 assessments were collected. NZIP AI assistant achieved the highest trustworthiness score (mean 1.21) but ranked lowest on clarity (1.53), user-friendliness (1.46), and relevance (1.65). General-purpose LLMs demonstrated the inverse pattern. Content analysis revealed that the NZIP AI assistant provided source citations in 100% of responses, compared with 0-3% for general-purpose LLMs; included explicit safety disclaimers in 85% of responses, compared with 5-18%; and declined 10.4% of queries outside its knowledge base.
CONCLUSIONS: National health information portals can integrate generative AI with RAG architecture to deliver verified, source-grounded health information, though trade-offs in user-oriented dimensions warrant consideration in future implementations. Still, this expert evaluation assessed response characteristics rather than direct effects on health literacy, awareness, or health outcomes.