科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on image processing : a publication of the IEEE Signal Processing Society2026-08-27

Broadcast-and-Mixing Transformer for 3D Semantic Segmentation.

Shuai Li, Yalong Bai, Wei Zhang, Hongjun Li, Qing Yang

原始摘要(英文原文)· Original abstract
Transformers have shown great promise in various point cloud comprehension tasks, but still face challenges due to the quadratic computational and memory cost when dealing with large-scale 3D point clouds. Many recent studies focus on reducing these costs and improving model performance by solely applying restricted local attention but overlook the coarse-grained global structural information, which is also crucial to 3D semantic segmentation. In this paper, we propose a novel Broadcast-and-Mixing Transformer model for 3D semantic segmentation. Leveraging the joint utilization of global, regional, and local structures within the point cloud, our approach first broadcasts the global representations learned by a lightweight voxel set attention to the regional level and then mixes them with local point features using a unique voxel-point self-attention mechanism. The model enables effective information exchange across different granularity levels, encompassing global-regional-local interactions, and controlling the overall computational complexity without a substantial increase after incorporating global information. Extensive experiments on large-scale indoor and outdoor datasets demonstrate the effectiveness of our proposed method, surpassing hybrid-input approaches and matching global-attention baselines with significantly lower memory cost.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Broadcast-and-Mixing Transformer for 3D Semantic Segmentation. — 科研速览 Science Skim