Rajat Gupta, Charu Gupta, Nitasha Rathore, Gargi Mishra
Intelligent surveillance systems require video anomaly detection methods that operate reliably under realworld conditions rather than controlled benchmark settings. This paper presents a deployment-orientedhybrid CNN–LSTM–MIL framework that integrates spatio–temporal feature learning, weakly supervisedanomaly scoring, and reconstruction-based regularity modeling to address the practical challenges oflarge-scale video surveillance. The proposed framework is evaluated on widely used benchmark datasets,including UCF-Crime, CUHK Avenue, ShanghaiTech, and UMN, as well as on diverse real-world CCTVfootage captured from urban streets, shopping malls, traffic intersections, and railway stations.Experimental results demonstrate competitive detection performance, achieving AUC scores of 85.9% onUCF-Crime and 91.3% on CUHK Avenue, while maintaining near real-time inference speeds of 28–50frames per second on GPU and edge platforms through deployment-oriented optimizations such aspruning and quantization. Additional evaluation on real-world surveillance data shows reduced falsealarm rates and stable detection performance under challenging conditions, including illuminationvariations, background clutter, occlusions, and varying crowd densities. By jointly analyzing detectionaccuracy, computational efficiency, and deployment feasibility, this work bridges the gap betweenbenchmark-oriented research and practical intelligent surveillance deployment for public safety andtraffic monitoring applications.