科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-08-28· cs.RO

CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

Zhenqi Wu, Yuanjie Lu, Yisheng Zhang, Miao Yu, Xuesu Xiao, Jaejeong Shin, Xiaomin Lin

原始摘要(英文原文)· Original abstract
Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments — 科研速览 Science Skim