科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cell reports. Medicine2026-08-11

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models.

Wuyang Lan, Siqi Zhang, Wenzheng Wang, Ke Hu, Tianrun Gao, Lei Shi, Zongbo Han, Yanjun Chen, Hao Zhang, Song Wu, Xiaohong Liu, Guangyu Wang

原始摘要(英文原文)· Original abstract
Biomedical information extraction (BIE) is fundamental for transforming unstructured biomedical text into structured, computable knowledge, yet the effectiveness of large language models (LLMs) remains limited by dataset heterogeneity and lack of unified benchmarks. We present InfoFlowEX, a unified framework for generalizable biomedical knowledge extraction with LLMs. InfoFlowEX incorporates an automated data integration pipeline using ontology-guided alignment to construct BIE-Corpus, a large-scale multi-domain benchmark unifying 40 public datasets for named entity recognition and relation extraction. We further introduce a task-conditioned schema instruction tuning strategy encoding 28 biomedical entity and relation types into a schema codebase, enabling LLMs to align heterogeneous annotations and generalize across settings. Finally, we evaluated InfoFlowEX in diverse applications, including evidence retrieval for question-answering, clinical diagnosis from electronic health records, and knowledge graph expansion. Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models. — 科研速览 Science Skim