Guanjie Wang, Lanxin Chen, Haoyang Bai, Zixiang Yi, Huaiyu Li, Dongxu Zhang
Reliable visual identification of household gas facilities is important for safety inspection, yet images acquired in real indoor environments are frequently affected by cluttered textures, shadows, metallic reflections, stains, dust, motion blur, occlusion, and viewpoint variation. In RT-DETR, these conditions can weaken cross-scale structural cues, reduce the ranking of small-object candidates, and amplify localization errors under strict IoU criteria. We hypothesize that coordinated intervention at feature fusion, query allocation, and box regression can alleviate this coupled failure process without enlarging the decoder-query budget. Accordingly, DRQ-RTDETR integrates degradation-aware detail recovery, small-object-guided query selection, and scale-adaptive geometric refinement. Experiments on a real household gas facility dataset containing 21,813 images and 47,169 instances across eight safety-related categories show that DRQ-RTDETR improves mAP from 0.6168 to 0.6576, mAP50 from 0.7909 to 0.8124, mAP75 from 0.6702 to 0.7136, and mAPsmall from 0.4639 to 0.5247 relative to RT-DETR. The larger gains in mAPsmall and mAP75 indicate that the proposed coordination is particularly effective for weak-response compact components and boundary-sensitive localization in degraded household scenes.