科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on image processing : a publication of the IEEE Signal Processing Society2026-08-07

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving.

Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou, Dingyuan Zhang, Hongwei Xie, Bing Wang, Xiang Bai

原始摘要(英文原文)· Original abstract
Large Vision-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on partial objects within scenes and simple question-answer pair annotations, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D spatial understanding and instruction following. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image-language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to provide structured geometric priors for language-conditioned 3D grounding. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a notable 9.86% improvement on the 3D visual grounding task. The dataset and code will be made available.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving. — 科研速览 Science Skim