科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ PloS one2026-01-01

Adaptive token division-based transformer for visual object tracking.

Dan Tian, Dongxin Liu, Xiao Wang

原始摘要(英文原文)· Original abstract
In visual object tracking, one-stream trackers typically use all search tokens to interact with templates across all encoder layers. However, the search area usually contains a lot of interference information, such as distractors with similar appearance to the tracking object, which will cause the distractors in the search area to be misjudged as interactive objects, establish wrong cross-relation modeling, and reduce the accuracy of tracking. To alleviate this issue, this paper proposes a transformer-based visual object tracking framework with adaptive token division. First, our tracking framework is a simple encoder-decoder structure without any post-processing. Second, we propose an adaptive token division module, which enables search tokens and template tokens to perform the most appropriate cross-relationship modeling, and improves the model 's ability to distinguish between object and background. At the same time, we introduce an attention masking strategy and Gumbel-Softmax technique. The strategy enables efficient and parallel attention calculation between different categories of tokens, and the technique facilitates the end-to-end optimization of the division module. Finally, we conduct tests on six tracking benchmarks, and the experimental results prove the effectiveness of our method.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Adaptive token division-based transformer for visual object tracking. — 科研速览 Science Skim