Deepak Ghimire, Donghoon Kim, Yeonho Jo, Eunhee Lee, Sunghwan Jeong, Byoungjun Kim
Generating early and reliable fire/smoke alarms from real-world video surveillance remains challenging because visually ambiguous patterns such as sunlight, reflections, clouds, mist, steam, and illumination changes often trigger unstable frame-level false alarms. This paper presents the Spatio-Temporal Verification Network for Fire and Smoke Alarm Validation (STV-FSANet), a detect-track-verify framework that first localizes candidate fire/smoke regions, associates them into temporal tracks, and then verifies the observed sequence of each active track online as fire, smoke, or false fire/smoke. The verifier combines a primary appearance stream from cropped candidate regions with lightweight geometric cues derived from bounding-box position, scale, motion, and short-term fluctuation. A dual-branch GRU models long-term track history, while recent temporal pooling emphasizes newly observed evidence for streaming decisions. To support temporal learning, we construct the Fire-Smoke Alarm Verification (FSAV) Tracklet Dataset from 1347 source videos, yielding 58,733 parent tracks and 2.06 million annotated track frames. The best matched-context STV-FSANet achieves 97.09% test accuracy and 95.81% macro-F1, and both shorter-context models are within 0.5 percentage points of their final prefix metrics by 1.0 s. The results indicate that compact temporal evidence is more useful than simply accumulating longer 96-frame histories. The proposed model also rejects false fire/smoke tracklets with 90.2% recall, demonstrating the value of explicit hard-negative modeling, while the decoupled design allows the verifier to be reused with future detector backbones. TensorRT FP16 deployment reaches 109.55 frames/s on an RTX 4060 Laptop GPU and 39.51 frames/s on a Jetson AGX Orin DevKit.