K. Shi, G.-W. Zhang, C. Tao, Z. Wang, S. K. Zhang, H. Tao, L. I. Zhang
Quantitative analysis of animal behavior is fundamental to neuroscience and ethology but remains constrained by the scalability, subjectivity, and limited reproducibility of manual annotation. Many automated approaches infer behavior through intermediate representations such as pose trajectories, whereas direct video modeling retains motion, appearance, and context, but its accuracy and efficiency remain incompletely established. Here we introduce V-TRACE (Video-based Temporal Recognition and Annotation of Continuous Ethograms of Animal Behavior), an end-to-end platform with a graphical user interface for behavioral detection and annotation. V-TRACE combines transformer-based video encoders with a multi-scale temporal module to model behavioral dynamics across a broad range of durations. Its modular design supports interchangeable encoder architectures within a common detection pipeline to produce frame-resolved behavioral identities and derive temporal boundaries from continuous recordings. Across datasets spanning species and experimental contexts, V-TRACE demonstrates high accuracy and high-throughput inference, enabling scalable, efficient, and context-aware animal behavior analysis directly from video.