科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Communications Medicine2025-10-15· Computer science

Introducing mCODEGPT as a zero-shot information extraction from clinical free text data tool for cancer research

Kai Zhang, Tongtong Huang, Bradley Malin, Travis Osterman, Qi Long, Xiaoqian Jiang

原始摘要(英文原文)· Original abstract
The vast amount of natural language clinical notes about patients with cancer presents a challenge for efficient information extraction, standardization, and structuring. Traditional NLP methods require extensive annotation by domain experts for each type of named entity and necessitate model training, highlighting the need for an efficient and accurate extraction method. This study introduces a tool based on the Large Language Model (LLM) for zero-shot information extraction from cancer-related clinical notes into structured data aligned with the minimal Common Oncology Data Elements (mCODE™) structure. We utilize the zero-shot learning capabilities of LLMs for information extraction, eliminating the need for data annotated by domain experts for training. Our methodology employs advanced hierarchical prompt engineering strategies to overcome common LLM limitations like token hallucination and accuracy issues. We tested the approach on 1,000 synthetic clinical notes representing various cancer types, comparing its performance to a traditional single-step prompting method. Our hierarchical prompt engineering strategy (accuracy = 94%, misidentification, and misplacement rate = 5%) outperforms the traditional prompt strategy (accuracy = 87%, misidentification, and misplacement rate = 10%) in information extraction. By unifying staging systems (e.g., TNM, FIGO) and specific stage details (e.g., Stage II) into a standardized framework, our approach achieves improved accuracy in extracting cancer stage information. Our approach demonstrates that LLMs, when guided by structured prompting, can accurately extract complex clinical information without the need for expert-labeled data. This method has the potential to harness unstructured data for advancing cancer research. Clinical notes about cancer patients contain a lot of valuable information, but they are often written in free text, making it hard for computers to use them in research. This study develops a new framework that uses large language models (LLMs) to automatically extract important details from these notes without needing manual labeling by experts. We introduce two advanced prompting techniques—BFOP and 2POP—that hierarchically guide the LLMs step-by-step through the information extraction process. We test BFOP and 2POP on 1,000 synthetic cancer notes, achieving high accuracy and low error rates. Our approaches could help researchers better understand cancer and make more informed clinical decision making by turning hard-to-read notes into structured, standardized data for analysis. Zhang, Huang et al. introduce mCODEGPT, a zero-shot framework that extracts named entities as structured data from clinical notes of cancer patients using Large Language Models. Their hierarchical prompt engineering approach enables accuracy information extraction without relying on expert-labelled training data.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Introducing mCODEGPT as a zero-shot information extraction from clinical free text data tool for cancer research — 科研速览 Science Skim