科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Sensors (Basel, Switzerland)2026-07-28

GazeHRNet: Head-Centric Spatial Encoding and Gaze-Aware Feature Interaction for Gaze Target Detection.

Tianxiang Nan, Chenglizhao Chen, Xi Chen, Zhi Li, Xiangyu Wei, Xinyu Liu

原始摘要(英文原文)· Original abstract
Gaze target detection requires understanding where a person is looking by jointly reasoning about the gazer and the surrounding scene. While recent methods have benefited from powerful pretrained visual backbones, they often treat gaze prediction as a generic localization problem and overlook a key property of the task: the target should be interpreted in relation to the person's head. This limits their ability to model direction, distance, and head-scene dependencies in a unified manner. We propose GazeHRNet, a head-centric reasoning framework for RGB-based gaze target detection. Instead of relying on absolute image coordinates or auxiliary geometric inputs, GazeHRNet represents the scene from the gazer's perspective through Head-Centric Polar Encoding and organizes visual features by their spatial relevance to the head via Head-Aware Attention Routing. It further combines coarse spatial reasoning with fine-grained anisotropic heatmap prediction, enabling reliable target localization under cluttered scenes and varying head positions. Experiments on GazeFollow and VideoAttentionTarget show that GazeHRNet achieves 0.952 and 0.929 AUC with L2 distances of 0.102 and 0.103, respectively, using only RGB input and 3 M trainable parameters. Cross-dataset evaluation further demonstrates improved robustness and generalization across different scenes and subject distributions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

GazeHRNet: Head-Centric Spatial Encoding and Gaze-Aware Feature Interaction for Gaze Target Detection. — 科研速览 Science Skim