Yaoxian Yang, Guipeng Lan, Shuai Xiao, Jiabao Wen, Jiachen Yang, Wen Lu, Baihua Li, Qinggang Meng, Xinbo Gao
These findings indicate that MoLiegroup supplies a practical, architecture-agnostic geometric inductive bias that improves expressivity and generalization across a wide range of vision tasks.
Deep visual models typically achieve robustness to geometric transformations through extensive data augmentation or increased model capacity, yet these empirical strategies do not guarantee explicitly equivariant or structurally constrained representations. While group-equivariant CNNs provide principled mechanisms for embedding symmetry, existing methods are commonly specialized to individual groups or struggle to balance expressivity when facing compound transformations. To address these limitations, we introduce MoLiegroup, a framework that embeds multiple Lie-group priors as specialized kernel experts and adaptively fuses them via a geometry-aware gating mechanism to learn equivariance-inspired geometric features. Each expert is parameterized from a spectral and representation-theoretic perspective: angular harmonic expansions encode rotation components, spatial harmonics model translations, and log-radial decompositions cover scale effects. An input-dependent router produces expert weights, and a load-balancing regularizer encourages uniform expert utilization, enabling the operator to interpolate between rotation-, translation-, and scale-equivariant responses at each spatial location. We validate MoLiegroup on both discriminative and generative tasks. These findings indicate that MoLiegroup supplies a practical, architecture-agnostic geometric inductive bias that improves expressivity and generalization across a wide range of vision tasks.