Yunjiao Zhou, Jianfei Yang, Han Zou, Lihua Xie
The rapid expansion of the Internet of Things (IoT) has introduced new challenges in Human Activity Recognition (HAR), particularly in dynamic environments where new and unforeseen activities emerge. Traditional HAR models, relying on predefined labels, struggle to adapt to these scenarios, highlighting the need for zero-shot learning (ZSL) approaches that can generalize beyond fixed training categories. Recent advances in large language models (LLMs) have demonstrated remarkable zero-shot capability in textual and visual domains. However, extending this ability to IoT sensors is substantially more challenging due to their heterogeneous modalities, diverse data structures, and limited semantic annotations. In this paper, we propose TENT (IoT-sEnsorslanguage alignmEnt pre-Training), a novel framework that constructs a unified sensor-language semantic space for zero-shot HAR. Instead of aligning each sensor individually to text, TENT jointly aligns multiple heterogeneous modalities with language, treating them as peers rather than anchors. This balanced multi-modal alignment allows sensors to mutually regularize one another while being grounded in linguistic semantics, transforming heterogeneity from a barrier into a strength. To further enrich the semantic space, TENT incorporates detailed activity descriptions and learnable prompts, enhancing adaptability to unseen activities. Extensive experiments across datasets and evaluation protocols demonstrate that TENT not only achieves robust recognition of both seen and unseen activities but also significantly outperforms existing vision-language and sensor-language baselines, surpassing them by over 20% on zero-shot HAR tasks. These results establish TENT as a new paradigm for generalizable IoT representation learning.