科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-09-21

Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four Approaches.

Doko Bandur, Milos Bandur, Branimir Jaksic

原始摘要(英文原文)· Original abstract
Vision Transformers (ViTs) rely on positional encoding (PE) because self-attention has no native notion of token order or image-grid location, yet the information-theoretic properties of different PE strategies and their downstream consequences for model behaviour remain insufficiently characterised. We present a systematic comparison of four PE approaches-Learned, Sinusoidal, Rotary Position Embedding (RoPE), and a 1D-ALiBi-style linear-bias variant-together with a targeted 2D-ALiBi-style diagnostic intervention motivated by a raster-distance mismatch diagnosis. ViT-Base models are trained on three low-to-intermediate data-regime datasets (CIFAR-100, TinyImageNet, and ImageNet-100), using a primary cross-method matrix and a canonical paired protocol for the targeted 2D-vs-1D ALiBi-style comparison. We apply a two-track diagnostic suite: embedding-space analyses (per-dimension variance, entropy, PCA, and probes) for additive PE, and attention-space intrinsic analyses (bias-tensor rank and entropy, slope schedule, and RoPE short-wavelength band counting) for attention-space PE, combined with attention-position mutual information (MI), noise ablation, and PE removal. The results reveal four qualitatively distinct encoding regimes: Learned PE is "quiet and ubiquitous," Sinusoidal PE "loud and structured," RoPE attains the highest accuracy with a graceful mid-depth MI decay, and the ALiBi-style variant exhibits a persistent attention-space bias. Direct attention-space analysis shows that the 1D-ALiBi-style bias tensor has full intrinsic rank. PE removal reveals a ${\sim }10\times$ dependency spectrum-from Sinusoidal collapse to near-chance to Learned PE retaining most of its accuracy-indicating extensive implicit positional learning. Linear probes show that Sinusoidal PE gives $0\%$ held-out column accuracy-a protocol-level generalisation failure for the modulo-column label rather than absence of column information-while Learned PE partially recovers 2D structure. Identifying the ALiBi-style raster-scan distance $|i{-}j|$ as mismatched to the 2D geometry of image patch grids, we introduce 2D-ALiBi-style, which replaces it with the 2D Euclidean patch-grid distance. On the canonical CIFAR-100 paired cohort ($n{=}12$ seeds), fixed-slope 2D-ALiBi-style yields a statistically reliable $+0.45$ pp accuracy improvement ($p{=}0.013$), is directionally positive on TinyImageNet, and shows no statistically reliable difference on ImageNet-100. It also lowers CLS-excluded patch-only attention-position MI. The mean-magnitude-matched 2D arm reverses this MI effect and reaches $70.28\pm 0.35\%$, the highest accuracy among the three canonical ALiBi-style arms; a secondary within-2D paired contrast ($+1.94$ pp, $12/12$ positive) establishes bias magnitude as a substantial lever without supporting a magnitude-rather-than-geometry interpretation. We therefore frame the contribution as a controlled characterisation plus a modest, statistically supported intervention, not a state-of-the-art PE method. Overall, PE choice affects accuracy, robustness, and attention-space information flow beyond what accuracy alone captures.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Information-Theoretic Analysis of Positional Encoding Strategies in Vision Transformers: A Comparative Study of Four Approaches. — 科研速览 Science Skim