科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Internet of Things Journal2026-05-19· Computer science

End-to-End 3-D Spatiotemporal Perception With Multimodal Fusion and V2X Collaboration

Zhenwei Yang, Yibo Ai, W Y Zhang

原始摘要(英文原文)· Original abstract
Multiview cooperative perception and multimodal fusion are essential for reliable 3D spatiotemporal understanding in autonomous driving, especially in cases with occlusions, limited viewpoints, and communication delays in vehicle-to-everything (V2X) scenarios. In this paper, XET-V2X, a multimodal fused end-to-end tracking framework for V2X collaboration that unifies multiview multimodal sensing within a shared spatiotemporal representation, is proposed. To efficiently align heterogeneous viewpoints and modalities, XET-V2X introduces a dual-layer spatial cross-attention module based on multiscale deformable attention. Multiview image features are aggregated to enhance semantic consistency, followed by point cloud fusion guided by the updated spatial queries, enabling effective cross-modal interaction while reducing computational overhead. Experiments based on the real-world V2X-Seq-SPD dataset and the simulated V2X-Sim-V2V and V2X-Sim-V2I benchmarks demonstrate consistent improvements in detection and tracking performance under varying communication delays, with XET-V2X achieving up to 15–20% relative gains in mAP and AMOTA over single-view or single-modal baselines, while also outperforming representative tracking-by-detection cooperative perception methods.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

End-to-End 3-D Spatiotemporal Perception With Multimodal Fusion and V2X Collaboration — 科研速览 Science Skim