科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-01· cs.CV

Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

Teresa DiMeola, Charles Walter, Hong Xiao

原始摘要(英文原文)· Original abstract
Global welfare often depends on the correct interpretation of aerial and satellite imagery. Acting on such imagery (mapping flooded ground, crop extent, or damaged infrastructure) demands pixel-level segmentation to ensure perfect class localization. Pretrained general foundation models, when applied directly, often miss important features and cannot always find all the classes belonging to a given scene, overlooking smaller objects that matter most. We use a single consumer-grade GPU running a vision-language model (VLM) to supply this missing guidance, improving segmentation while producing structured, auditable evidence that drives the result and can be inspected on its own. We fuse three approaches: the frozen foundation model that labels every pixel, and two queries to a VLM, one to choose the classes that matter, and one to locate the small objects the base model misses. Evaluating across four aerial datasets, we see consistent gains at each stage where the base model is competent.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation — 科研速览 Science Skim