Juan José Cabrera, Antonio Santo, Arturo Gil, Carlos Viegas, Luis Payá
This paper presents MinkUNeXt, an effective and efficient architecture for place-recognition from point clouds entirely based on the new 3D MinkNeXt Block, a residual block composed of 3D sparse convolutions that follows the philosophy established by recent Transformers but purely using simple 3D convolutions. Feature extraction is performed at different scales by a U-Net encoder–decoder network and the feature aggregation of those features into a single descriptor is carried out by a Generalized Mean Pooling (GeM). The proposed architecture demonstrates that it is possible to surpass the current state-of-the-art by only relying on conventional 3D sparse convolutions without making use of more complex and sophisticated proposals such as Transformers, Attention-Layers or Deformable Convolutions. A thorough assessment of the proposal has been carried out using the Oxford RobotCar, the In-house, the KITTI and the USyd datasets. As a result, MinkUNeXt proves to outperform other methods in the state-of-the-art. The implementation is publicly available at https://juanjo-cabrera.github.io/projects-MinkUNeXt/ . • MinkUNeXt: The first U-Net architecture devised for point cloud embedding and place recognition. • 3D MinkNeXt Block: A novel residual block with 3D sparse convolutions outperforming ResNet. • A detailed ablation study validating each architectural design choice and its impact. • An efficient approach achieving superior results without complex attention mechanisms. • A comprehensive evaluation showing state-of-the-art results on multiple benchmark datasets.