Haodong Liu, Linjing Wei, Pengyu Hou, Yi Hu, Yifei Ding, Yangbin Li
On the public Lincoln Beet benchmark, ASSET-DETR attains 74.2% mAP50 and 51.7% mAP50-95, outperforming twelve representative detectors under a unified protocol, and improves on RT-DETR-R18 by 7.8 and 7.7 percentage points at 21.34 M parameters and 100 FPS. On external datasets, it retains about 56% of its source-domain accuracy, and on OD-SugarBeets, few-shot fine-tuning with a small amount of target-domain annotation restores performance to near source-domain levels.
INTRODUCTION: At the sugar beet seedling stage, crops and weeds look alike, vary widely in scale, and appear against heterogeneous backgrounds, so a detector must judge how reliable each location and scale is, not merely extract multi-scale features.
METHODS: Building on RT-DETR, we propose ASSET-DETR, a relation-aware detector rebuilt at three levels: a sparse-dense attention encoder (ASSET), an attention-gated backbone (CAG-Backbone), and a hypergraph-based neck (HyperFPN). Its Hypergraph-Conditioned Multi-Feature Fusion Module (HC-MFM) conditions the fusion weights on group-level relational semantics, turning multi-scale fusion from aggregation by local statistics into feature arbitration.
RESULTS: On the public Lincoln Beet benchmark, ASSET-DETR attains 74.2% mAP50 and 51.7% mAP50-95, outperforming twelve representative detectors under a unified protocol, and improves on RT-DETR-R18 by 7.8 and 7.7 percentage points at 21.34 M parameters and 100 FPS. On external datasets, it retains about 56% of its source-domain accuracy, and on OD-SugarBeets, few-shot fine-tuning with a small amount of target-domain annotation restores performance to near source-domain levels.
DISCUSSION: ASSET-DETR is therefore a parameter-efficient candidate for object-level perception in variable-rate spraying and precision weeding, and the results underline the importance of cross-domain degradation and few-shot adaptation for this task.