Xinlong Wang, Zongxuan Li, Shuping Tao, Lin Li, Yu Zhao, Zhiyuan Gu, Zhen Wen, Jiangtao Yu
Off-axis reflective optical systems offer irreplaceable advantages in space exploration and high-resolution imaging owing to their unobscured and compact design. However, their inherent asymmetry induces severe aberration coupling between rigid-body misalignments and residual surface figure errors, causing traditional sensitivity matrix methods and conventional deep learning approaches to suffer from unstable convergence or complete failure. This paper proposes a field-aware physics-informed transformer, termed FA-PIT, that reformulates multi-field Zernike coefficients as structured field tokens to preserve the spatial topology of the full field of view. A multi-head self-attention mechanism autonomously learns cross-field aberration coupling patterns and tracks global nodal drift induced by asymmetric misalignments. A physics-informed composite loss function integrates a sensitivity-matrix-based full-field consistency constraint, a uniformity penalty, and a worst-field diffraction-limit safety term to suppress physically unreasonable compensation commands. Extensive Monte Carlo simulations under surface figure errors from λ/50 to λ/20 and initial misalignments of ±0.1 mm/0.1∘ demonstrate that FA-PIT achieves success rates of 99%, 96%, and 69% with final worst-field RMS values of 0.029λ, 0.041λ, and 0.064λ, substantially outperforming five representative baselines including FSAHNet, ResMLP, FCNN, the second-order matrix method, and the NAT-based analytical method. Under extrapolated out-of-distribution conditions with misalignment ranges expanded to ±0.5 mm, FA-PIT retains a success rate of 73% to 90% while competing methods drop to 20% to 65%. Under Zernike measurement noise up to σ = 0.005λ, FA-PIT maintains success rates of 58% to 85%. Cross-system transfer experiments on an independent off-axis architecture confirm a 93% success rate. Ablation experiments verify that field tokenization, multi-head self-attention, learnable field-position embedding, and the physics-informed composite loss all contribute essentially to the robustness of FA-PIT, and that the performance gain cannot be attributed merely to parameter scale or training strategy. These results establish FA-PIT as an effective, physically interpretable, and robust framework for the automated alignment of off-axis reflective optical systems.