Zekun Qian, Wei Feng, Ruize Han, Junhui Hou
Multi-object tracking (MOT) has traditionally focused on a few specific categories, thereby restricting its applicability to real-world scenarios involving diverse objects. Open-vocabulary multi-object tracking (OVMOT) addresses this limitation by enabling tracking of arbitrary categories, including novel objects unseen during training. However, current progress is constrained by two critical challenges: the lack of continuously annotated video data for model training, and the lack of a customized OVMOT framework to synergistically handle the sub-tasks, i.e., detection (including localization and classification) and association. We address the data bottleneck by constructing C-TAO, the first continuously annotated training set for OVMOT, which increases annotation density by 26× over the original TAO and captures smooth motion dynamics and intermediate object states. For the framework bottleneck, we propose COVTrack++, a synergistic framework that achieves a bidirectional reciprocal mechanism between detection and association. This is realized across three modules: (1) Multi-cue adaptive fusion module dynamically balances the appearance, motion, and semantic cues for association feature learning; (2) Multi-granularity hierarchical aggregation module further exploits hierarchical spatial relationships in dense detections, where visible child nodes (e.g., object parts) assist occluded parent objects (e.g., whole body) for association feature enhancement; and (3) Temporal confidence propagation module recovers flickering detections through high-confidence tracked objects, boosting low-confidence candidates across frames and creating a chain reaction that stabilizes trajectories. Extensive experiments on TAO demonstrate the state-of-the-art performance of COVTrack++, with novel TETA reaching 35.4% and 30.5% on validation and test sets, respectively, improving novel AssocA by 4.8% and novel LocA by 5.7% over previous methods, and show strong zero-shot generalization on BDD100K. These results validate the effectiveness of the C-TAO dataset and the robustness of COVTrack++.