科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JACC. Asia2026-07-30

Comprehensive Echocardiography Interpretation Using Video and Multiview Vision-Language AI.

Ryo Takizawa, Chiemi Yamazaki, Satoshi Kodera, Tempei Kabayama, Shun Kitamura, Risa Kishikawa, Ryo Matsuoka, Megumi Hirokawa, Junichi Ishida, Koki Nakanishi, Hiroyuki Morita, Norihiko Takeda

一句话结论 · In one sentence

A multiview video-language framework improved report retrieval compared with conventional image-based approaches. These findings support the utility of video-based, multiview representation learning for echocardiographic report retrieval.

原始摘要(英文原文)· Original abstract
BACKGROUND: Echocardiography is essential for assessing cardiac structure and function, yet accurate interpretation requires specialized expertise, creating challenges in emergency care and in regions with limited access to experienced echocardiographers. Although artificial intelligence-based interpretation is increasingly studied, most existing models rely on still images or single views and do not reflect the video-based, multiview integration used in clinical practice. OBJECTIVES: The authors aim to develop and evaluate a multiview video-language model that integrates cardiac motion and multiple standard echocardiographic views. METHODS: We trained a vision-language model on 577,061 transthoracic echocardiography videos paired with Japanese clinical reports from 46,852 examinations acquired at a single tertiary center (2015-2023). For each examination, features from 5 standard views (parasternal long-axis and short-axis; apical 2-, 3-, and 4-chamber) were aggregated. Performance was assessed in a retrieval task selecting the correct report from 16,436 candidate reports in the test cohort based on video input; the primary metric was correct match retrieval probability within the top 10 candidates (ie, R@10). RESULTS: The image-based baseline achieved an R@10 of 1.3%. Using video input increased R@10 to 4.2%. Integrating 5 views further improved R@10 to 6.0% (95% CI: 5.5%-6.4%). Gains were greatest for findings that depended on temporal dynamics and multiview assessment, including left ventricular wall motion abnormality and left ventricular dilation. CONCLUSIONS: A multiview video-language framework improved report retrieval compared with conventional image-based approaches. These findings support the utility of video-based, multiview representation learning for echocardiographic report retrieval.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comprehensive Echocardiography Interpretation Using Video and Multiview Vision-Language AI. — 科研速览 Science Skim