科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Multimodal Technologies and Interaction2026-07-31· Debiasing

Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering

Yali Fan, 黄刚玉, Qiwen Lu, Shengbo Chen

原始摘要(英文原文)· Original abstract
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and generalization. Existing debiasing techniques mainly depend on post hoc evaluation or architectural modifications, while recent prompt-learning-based methods reveal new opportunities for aligning downstream tasks with pretrained models. In this work, we propose a unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism. The generative prompt component reformulates VQA as a cloze-style masked prediction problem, leveraging pretrained language priors to improve semantic grounding. Meanwhile, the fuzzing-based module actively constructs unexpected test samples during training and employs a reflection mechanism to correct biased predictions in-loop, yielding inference-time robustness without additional test-time components. Extensive experiments on VQA-v2, VQA-CP, VQA-CE, GQA-OOD, and VQA-VS demonstrate that the proposed framework significantly improves both in-distribution (ID) accuracy and out-of-distribution (OOD) robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering — 科研速览 Science Skim