科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Expert Systems with Applications2026-05-22· Hallucinating

Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation

Musarrat Zeba, Abdullah‐Al Mamun, Kishoar Jahan Tithee, Debopom Sutradhar, Mohaimenul Azam Khan Raiaan, Saddam Hossain Mukta, Reem E. Mohamed, Md. Rafiqul Islam, Yakub Sebastian, Mukhtar Hussain, Sami Azam

原始摘要(英文原文)· Original abstract
In healthcare, it is essential for any Large Language Model (LLM)-generated output to be reliable and accurate, particularly in cases involving decision-making and patient safety. However, the outputs are often unreliable in such critical areas due to the risk of hallucinated outputs from the LLMs. To address this issue, we propose a fact-checking module that operates independently of any LLM, along with a domain-specific summarization model designed to minimize hallucination rates. Our model is fine-tuned using Low-Rank Adaptation (LoRA) on the MIMIC-III dataset and is paired with the fact-checking module, which uses numerical tests for correctness and logical checks at a granular level through discrete logic in natural language processing (NLP) to validate facts against electronic health records (EHRs). We trained the LLM on the full MIMIC-III dataset. For evaluation of the fact-checking module, we sampled 104 summaries, extracted them into 3786 propositions, and used these as facts. The fact-checking module achieves a precision of 0.8904, a recall of 0.8234, and an F1-score of 0.8556. Additionally, the LLM summary achieves a ROUGE-1 score of 0.5797 and a BERTScore of 0.9120 for summary quality.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation — 科研速览 Science Skim