Yeon Woo Cho, Jung Woo Cheon, Seok Bong Yoo
• Proposes a structure-aware monocular 3D detector based on semantic keypoints. • Recovers occluded object structure using a direction-guided autoencoder. • Uses dual-band depth features to mitigate depth ambiguity under occlusion. • Achieves reliable 3D perception for safety-critical transportation systems. Monocular 3D object detection provides a cost-effective alternative to multi-sensor setups, but estimating accurate depth from a single image remains inherently challenging. The difficulty becomes more pronounced when objects are partially visible. To address these challenges, we introduce skeleton keypoints. Semantic skeleton keypoints serve as interpretable structural cues that encode expert knowledge about object geometry. Our keypoint structural representation detects visible skeleton keypoints and reconstructs missing ones with an autoencoder reconstructor conditioned on coarse pose, providing a compact and shape-aware cue tolerant to occlusion and truncation. In parallel, depth features are estimated by a dual-band-guided depth predictor, which explicitly decomposes low- and high-frequency components to balance global context and fine details, thereby mitigating spectral bias and improving depth estimation for tiny and distant objects. These representations are then fused, and a scene-topology 3D detection head decodes class and box parameters while conditioning on neighborhood topology to enforce geometric consistency. Experiments show that our method provides more reliable 3D perception for decision-support systems in intelligent transportation by improving robustness under occlusion, demonstrating its effectiveness for autonomous driving perception systems.