科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-30· health informatics

Novel document-level measures and their performance for active learning uncertainty sampling: a use case for automatic CDSS ontology curation

S. Alluri, K. Komatineni, R. Goli, R. D. Boyce, N. Hubig, H. Min, Y. Gong, D. F. Sittig, D. Robinson, P. Biondich, A. Wright, C. Nohr, T. D. Law, A. Faxvaag, R. W. Gimbel, L. Rennert, X. Jing

原始摘要(英文原文)· Original abstract
Human-in-the-loop active learning (AL) can identify key phrases to facilitate biomedical ontology development, maintenance, and other curation tasks. However, determining which documents to annotate by humans is not straightforward. We explored new strategies to make the document selection process transparent, reproducible, and effective. Our AL pipeline modified a BiLSTM-CRF model using PubMed abstracts. We tested four novel document-level uncertainty aggregation strategies: KPSum, KPAvg, DocSum, and DocAvg, that operate over standard token-level uncertainty scores: Minimum Token Probability (MTP), Token Entropy (TE), and Margin. All strategies show significant improvement in early active learning cycles ({theta} to {theta}2) for recall and F1. The systematic evaluations show that KPSum (actual order) shows consistent improvement in both recall and F1. The weighted F1 ({beta} = 5, 10) provided complementary results to raw recall and F1 ({beta} = 1). Our work advances uncertainty sampling by introducing document-level uncertainty aggregation and shows promise in automating ontology curation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Novel document-level measures and their performance for active learning uncertainty sampling: a use case for automatic CDSS ontology curation — 科研速览 Science Skim