科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

Robust vision transformer adaptation for UAV imagery.

Jagan Murugesan, P Helen Vijitha, L Jai Vinita, A K Parvathy, P N Jeipratha

原始摘要(英文原文)· Original abstract
Vision Transformers (ViTs) perform well on clean aerial imagery but degrade sharply when deployed in post-earthquake UAV operations, where motion blur, dust haze, illumination variation, and sensor noise combine to produce what we term post-earthquake visual drift, a structured distributional shift that can render an otherwise capable model dangerously unreliable in the field. Retraining or conventional domain adaptation is not a realistic option under the latency, compute, and annotation constraints of active disaster response. This work presents a comprehensive empirical evaluation of lightweight label-free test-time adaptation (TTA) methods for improving the robustness of ViTs under such deployment conditions, systematically comparing consistency-based self-supervision (MEMO) and entropy minimization (TENT) across datasets with varying classification complexity. Consistent with existing lightweight TTA approaches, only the LayerNorm affine parameters are adapted during inference, while the backbone and attention weights remain frozen. On the UAV-TEBDE post-earthquake dataset, accuracy rises from 73.47% under severe drift (severity 0.3) to 89.18% using MEMO, corresponding to a recovery of 15.71 percentage points over the drifted baseline, without using any ground-truth labels or performing any retraining. On UAV-TEBDE, MEMO also achieved higher recovery than the SAR baseline evaluated in this study. Additional evaluation on AID (30-class aerial scene classification) and UC Merced (21-class land-use classification) shows consistent drift-induced degradation and adaptation-driven recovery across datasets of varying complexity. The cross-dataset evaluation further reveals that the relative effectiveness of lightweight TTA methods depends on the classification-space dimensionality. TENT-based entropy minimization outperforms consistency-based MEMO when the class count is large, while MEMO is the stronger choice in low-class settings. This interaction between method design and classification-space dimensionality has direct consequences for choosing TTA strategies in operational disaster assessment pipelines. Feature-space PCA analysis confirms that LayerNorm-only updates geometrically recalibrate internal representations, rather than simply correcting the model's output logits.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Robust vision transformer adaptation for UAV imagery. — 科研速览 Science Skim