科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Talanta2026-08-31

VL-ThyNet: An interpretable analytical framework integrating visual foundation models and vision-language reasoning for malignancy risk assessment of C-TIRADS 4 thyroid nodules.

Yao Chen, Xingyu Chen, Junyang Hu, Miduo Tan, Yingqi Wang, Libo Nie, Tong Wang

原始摘要(英文原文)· Original abstract
Accurate differentiation between benign and malignant C-TIRADS 4 thyroid nodules is particularly challenging because of their intermediate malignancy risk and overlapping sonographic features, yet it is essential for clinical decision-making. Herein, we propose VL-ThyNet, an interpretable analytical framework for malignancy risk assessment. VL-ThyNet integrates a DINOv2 vision foundation model with a ResNet adapter to jointly capture global semantic representations and local spatial-textural features from transverse and longitudinal views. In parallel, a vision-language model generated structured radiological descriptors together with confidence scores, providing clinically interpretable radiological information for subsequent feature fusion. Ablation experiments evaluated the effects of dual-view input, the DINOv2 encoder, and VLM-derived radiological features, showing that their contributions varied across model configurations. On the test set, VL-ThyNet achieved an accuracy of 84.8% and an AUC of 0.902, showing higher ACC and AUC than the evaluated conventional CNN- and Transformer-based models. By integrating Grad-CAM visual localization with structured VLM-based interpretation, VL-ThyNet provided clinically aligned diagnostic evidence. In a reader study, VL-ThyNet assistance was associated with improved diagnostic accuracy among radiologists and an 18.4% reduction in the average image interpretation time of the six radiologists, indicating its potential as an interpretable and efficient auxiliary analytical tool for thyroid nodule assessment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

VL-ThyNet: An interpretable analytical framework integrating visual foundation models and vision-language reasoning for malignancy risk assessment of C-TIRADS 4 thyroid nodules. — 科研速览 Science Skim