科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Sensors (Basel, Switzerland)2026-09-05

Benchmarking Zero-Shot Open-Vocabulary and Fine-Tuned Object Detectors for Underground Mine Personnel Detection.

Ellen Essien, Samuel Frimpong

原始摘要(英文原文)· Original abstract
Reliable personnel detection is critical for the safe deployment of autonomous haulage systems in underground mining, where challenging environmental conditions demand robust real-time perception. Existing research has focused primarily on fine-tuned convolutional detectors, while systematic comparisons with zero-shot vision-language models remain limited. This study presents a cross-paradigm benchmark comparing four zero-shot vision-language models (YOLO-World, Grounding DINO, OWL-ViT, and OWLv2) with four fine-tuned YOLO detectors (YOLOv8s, YOLOv9s, YOLO11s, and YOLO26s) using 31,396 real-world underground coal mine images. Detection performance was evaluated using precision, recall, F1-score, average precision, inference speed, and condition- and target scale-specific recall. The experimental results show that the fine-tuned detectors achieved F1-scores of 0.8449-0.8662 and AP50 values of 0.8791-0.9135, with YOLO26s achieving the strongest overall performance. In comparison, the zero-shot models achieved F1-scores of 0.3186-0.5484 and AP50 values of 0.2484-0.5108, with Grounding DINO performing best among the zero-shot models. Fine-tuned detectors also maintained substantially higher recall for occluded personnel and small apparent targets. These findings demonstrate a substantial performance advantage for domain-specific fine-tuning over the evaluated zero-shot approaches and establish a controlled cross-paradigm benchmark for comparing detection paradigms for underground personnel perception.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking Zero-Shot Open-Vocabulary and Fine-Tuned Object Detectors for Underground Mine Personnel Detection. — 科研速览 Science Skim