科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-09-16

Kalmanized Recursive Least Squares Attention for Streaming Video Understanding.

Matthew Korban, Peter Youngs, Scott T Acton

原始摘要(英文原文)· Original abstract
Streaming long-form video understanding requires online temporal reasoning over sequences far longer than typical training clips. Standard temporal self-attention is poorly matched to this setting: its cost grows quadratically with context length, it often degrades under context extrapolation, and it provides no native uncertainty estimate for temporally out-of-distribution events. We propose Kalmanized Recursive Least Squares Attention (KRLS-Attn), a drop-in replacement for temporal attention that recasts attention as online kernel ridge regression. KRLS-Attn maintains a fixed-size memory state-the posterior mean and covariance of a linear predictor-updated exactly via a recursive least-squares/Kalman recursion, eliminating explicit temporal attention matrices while preserving query-conditioned aggregation over the full history and keeping memory and per-step computation independent of stream length. The posterior covariance yields token-level predictive uncertainty, enabling adaptive token admission and confidence-weighted temporal pooling, while regularization and forgetting provide explicit control under temporal drift. We also derive the method's connections to kernel ridge regression and Bayesian filtering, including exact online updates, predictive uncertainty, and stability guarantees. Integrated into a factorized video Transformer, KRLS-Attn yields strong accuracy-efficiency trade-offs and improved context-length extrapolation on Kinetics-400, Something-Something V2, EPIC-KITCHENS-100, and Breakfast.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Kalmanized Recursive Least Squares Attention for Streaming Video Understanding. — 科研速览 Science Skim