Stefan Grushko, Aleš Vysocký, Jakub Chlebek
Hand localization in cluttered industrial environments remains challenging due to variations in appearance and the gap between synthetic and real-world data. Domain randomization addresses this “reality gap” by intentionally introducing randomized and unrealistic visual features in simulated scenes, encouraging neural networks to focus on essential domain-invariant cues. In this study, we applied domain randomization to generate a synthetic Red-Green-Blue–Depth (RGB-D) dataset for training multimodal instance segmentation models, with the aim of achieving color-agnostic hand localization in complex industrial settings. We introduce a new synthetic dataset tailored to various hand detection tasks and provide ready-to-use pretrained instance segmentation models. To enhance robustness in unstructured environments, the proposed approach employs multimodal inputs that combine color and depth information. To evaluate the contribution of each modality, we analyzed the individual and combined effects of color and depth on model performance. All evaluated models were trained exclusively on the proposed synthetic dataset. Despite the absence of real-world training data, the results demonstrate that our models outperform corresponding models trained on existing state-of-the-art datasets, achieving higher Average Precision and Probability-Based Detection Quality.