科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of imaging informatics in medicine2026-08-26

EndoVLM: A Vision-Language Assistant for Gastrointestinal Endoscopy.

Shuai Wang, Chengbo Liu, Tiejun Zhao, Bin Zhao

原始摘要(英文原文)· Original abstract
Gastrointestinal endoscopy generates extensive high-resolution video data, posing significant challenges for efficient and accurate computer-aided diagnosis of gastrointestinal diseases. To address this, we propose EndoVLM (Endoscopy Vision-Language Model), a specialized visual question-answering assistant for gastroenterology. EndoVLM introduces ConvNeXt as a hierarchical visual encoder to replace traditional ViTs (Vision Transformers), inherently compressing high-resolution gastrointestinal endoscopy images into information-dense features while reducing redundant token overhead. In addition, EndoVLM adopts a three-stage fine-tuning schedule that progressively aligns the visual-language projector, adapts the ConvNeXt backbone to endoscopic imagery, and improves instruction following in the language decoder. We position this design as an efficient adaptation strategy for high-resolution gastrointestinal endoscopy rather than as a new fine-tuning paradigm, with the aim of balancing token efficiency, domain adaptation, and deployment practicality in multimodal medical AI (artificial intelligence). Quantitatively, EndoVLM achieves 0.7326 ± 0.0038 ROUGE-1 (Recall-Oriented Understudy for Gisting Evaluation-1) and 0.5103 ± 0.0051 BLEU (Bilingual Evaluation Understudy) on Kvasir-VQA (Kvasir visual question answering), improving over the strongest public baseline by +0.0166 ROUGE-1 and +0.0323 BLEU, while the final three-stage model improves Gastrovision external validation from 0.4760 to 0.5799 Accuracy and from 0.1109 to 0.1268 macro-F1 (macro-averaged F1 score) over the representative two-stage baseline. The experimental section separates descriptive benchmark comparison from controlled follow-up analyses, including stage-wise convergence evidence, external validation on Gastrovision, and stage-wise training-strategy analysis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

EndoVLM: A Vision-Language Assistant for Gastrointestinal Endoscopy. — 科研速览 Science Skim