Til Spreuer, Andreas Günther, Raphael W Majeed
LLMs excel at extraction. They are suitable for information extraction of medications from clinical notes for use in research databases. However, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.
INTRODUCTION: Accurate clinical coding is fundamental to large-scale epidemiological studies, hospital billing, and the development of robust clinical decision support systems. Conventional methods for structured data extraction often rely on manual curation, which is prohibitively labor-intensive. Goal of this project is to determine whether current state-of-the-art open-weight LLM models are suitable for extraction of structured data from non-English (German) clinical notes.
METHODS: We anonymized 35 German doctor's notes of five patients from our hospital and developed one pipeline to extract and map medications and two for diagnoses. The latter compares a RAG based approach with an agentic AI. We ran these using three open-weight LLMs on a local GPU-PC.
RESULTS: The F1 scores for diagnoses do not exceed 0.12. If we instead consider mapping to the broad category, then the F1 score increases to 0.18. For medications, the F1 score is as high as 0.78 and even 0.95 if we consider trivial name extraction only.
DISCUSSION: For trivial name extraction of medications, every encountered mistake is explainable. Due to limitations in the nature of the task, it is infeasible to expect a perfect score of 1 in any of the coding scenarios. Further problems in LLM output and parsing are addressed.
CONCLUSION: LLMs excel at extraction. They are suitable for information extraction of medications from clinical notes for use in research databases. However, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.