Zhipeng Yin, Zichong Wang, Min Chen, Ian Stockwell, Xin Ning, Jun Liu, Wenbin Zhang
The application of question-answering (QA) systems in the medical domain has rapidly advanced, significantly improving patients' access to reliable health-related information. However, current approaches face notable challenges, including the difficulty in obtaining large-scale and unbiased medical datasets, significant privacy concerns, and inefficiencies due to manual dataset annotation. To address these issues, This study introduce a novel methodology leveraging publicly accessible online health forums to systematically build an unbiased, privacy-conscious QA dataset, and it employs Topic-guided Semantic Modeling (TGSM) for automated topic identification, enabling efficient and targeted annotation of relevant patient-generated content. Subsequently, this study propose a two-stage QA pipeline based on a Retriever-Reader architecture, which is further enhanced through fine-tuning state-of-the-art transformer-based models on the constructed domain-specific QA dataset. Experimental results demonstrate that our fine-tuned BioBERT significantly outperforms existing benchmarks, offering accurate patient-derived insights and providing a replicable framework for building efficient medical QA systems.