科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of imaging informatics in medicine2026-08-10

CONRep: Uncertainty-Aware Vision-Language Report Drafting Using Conformal Prediction.

Danial Elyassirad, Benyamin Gheiji, Mahsa Vatanparast, Amir Mahmoud Ahmadzadeh, Seyed Amir Asef Agah, Mana Moassefi, Meysam Tavakoli, Shahriar Faghani

原始摘要(英文原文)· Original abstract
The objective of this study is to quantify uncertainty in vision-language model (VLM)-based automated radiology report drafting (ARRD) to support trustworthy clinical deployment. In this study, we developed CONRep in two settings: label-based and sentence-based. In the label-based pipeline, we used the ChestX-Det10 dataset, comprising 3001 chest radiographs with radiologist annotations for ten thoracic pathologies. Two VLM paradigms were evaluated for classification, and conformal prediction (CP) was applied to stratify outputs into certain (high-confidence) and uncertain (low-confidence) subgroups at three significance levels (α = 0.05, 0.1, 0.2). Classification performance was compared between subgroups using the area under the receiver operating characteristic curve (AUROC). In the sentence-based pipeline, experiments were conducted on the Open-I chest X-ray dataset (3660 cases), using radiologist-written impression sections as ground-truth (GT). A decoder-only VLM generated impression texts, and CP was applied to categorize outputs as certain or uncertain/highly uncertain subgroups. Semantic alignment between generated and GT impressions was quantified using cosine similarity computed with a contrastive VLM and compared between subgroups. Additionally, an independent large language model (LLM) was used to score semantic agreement between generated and GT impressions and to judge clinical concordance, providing an evaluation independent of the scoring embedding. Across both pipelines, outputs classified as certain demonstrated significantly higher agreement with GT than uncertain outputs. In the label-based experiments, the certain subset achieved significantly higher AUROCs across multiple findings (P < 0.05). In the sentence-based experiments, certain cases showed significantly greater semantic similarity between generated and reference impressions (P < 0.001). This was corroborated by the independent LLM evaluator, which assigned significantly higher agreement scores and match rates to certain cases than to uncertain/highly uncertain cases (P < 0.001). CONRep provides a model-agnostic framework for uncertainty-aware ARRD using CP. By quantifying predictive uncertainty, CONRep enhances the transparency, reliability, and clinical usability of VLM-based ARRD systems, supporting safer integration into radiology workflows.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

CONRep: Uncertainty-Aware Vision-Language Report Drafting Using Conformal Prediction. — 科研速览 Science Skim