科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Image and Vision Computing2026-03-03· Computer science

All you need for object detection: From pixels, points, and prompts to Next-Gen fusion and multimodal LLMs/VLMs in autonomous vehicles

Sayed Pedram Haeri Boroujeni, Niloufar Mehrabi, Hazim Alzorgan, Mahlagha Fazeli, Abolfazl Razi

原始摘要(英文原文)· Original abstract
Autonomous Vehicles (AVs) are transforming the future of transportation through advances in intelligent perception, decision-making, and control systems. However, their success is tied to one core capability, reliable object detection in complex and multimodal environments. While recent breakthroughs in Computer Vision (CV) and Artificial Intelligence (AI) have driven remarkable progress, the field still faces a critical challenge as knowledge remains fragmented across multimodal perception, contextual reasoning, and cooperative intelligence. This survey bridges that gap by delivering a forward-looking analysis of object detection in AVs, emphasizing emerging paradigms such as Vision-Language Models (VLMs), Large Language Models (LLMs), and Generative AI rather than re-examining outdated techniques. We begin by systematically reviewing the fundamental spectrum of AV sensors (camera, ultrasonic, LiDAR, and Radar) and their fusion strategies, highlighting not only their capabilities and limitations in dynamic driving environments but also their potential to integrate with recent advances in LLM/VLM-driven perception frameworks. We also review autonomous vehicle simulators as a critical layer for safe development, scalable testing, and reproducible benchmarking of perception and detection pipelines before real-world deployment. Next, we introduce a structured categorization of AV datasets that moves beyond simple collections, positioning ego-vehicle, infrastructure-based, and cooperative datasets (e.g., V2V, V2I, V2X, I2I), followed by a cross-analysis of data structures and characteristics. Ultimately, we analyze cutting-edge detection methodologies, ranging from 2D and 3D pipelines to hybrid sensor fusion, with particular attention to emerging transformer-driven approaches powered by Vision Transformers (ViTs), Large and Small Language Models (SLMs), and VLMs. By synthesizing these perspectives, our survey delivers a clear roadmap of current capabilities, open challenges, and future opportunities, highlighting underexplored avenues such as multimodal reasoning, cooperative perception, and foundation-model integration. We aim to establish this work as a definitive reference for researchers, practitioners, and developers, fostering accelerated innovation toward safer and more intelligent autonomous driving systems. • Comprehensive review of state-of-the-art object detection in autonomous vehicles. • Analysis of latest AV sensors, fusion strategies, and multimodal perception systems. • Novel categorization and comparison of ego-vehicle, roadside, and CP datasets. • In-depth evaluation of 2D, 3D, fusion, and emerging LLM/VLM-based detection methods. • Highlights open challenges and potential advancements in AV perception research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

All you need for object detection: From pixels, points, and prompts to Next-Gen fusion and multimodal LLMs/VLMs in autonomous vehicles — 科研速览 Science Skim