Diyar Altinses, Andreas Schwung
Ensuring the stability and robustness of multimodal autoencoders is critical for their optimization and deployment in safety-critical industrial environments. This paper presents a rigorous analysis of Lipschitz properties in multimodal fusion, identifying theoretical vulnerabilities in standard summation and concatenation strategies. We derive explicit Lipschitz bounds for these methods, demonstrating their susceptibility to gradient instability and noise propagation as the number of modalities increases. To address this, we introduce a Lipschitz-regularized attention-based fusion mechanism that explicitly bounds gradient sensitivity through spectral normalization and dimension scaling. Empirical validation on four industrial robotic datasets, including the real-world RoboMNIST dataset, confirms our theoretical findings. On the RoboMNIST dataset, our method achieves an improvement of 54.7% in bimodal and 57.4% in trimodal reconstruction over standard attention, alongside a 20.5% gain in subordinated fault detection, while effectively regularizing the Lipschitz of the gradients.