科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural networks : the official journal of the International Neural Network Society2026-09-16

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

Xilang Huang, Seon Han Choi

原始摘要(英文原文)· Original abstract
Few-shot open-set recognition (FSOR) presents unique challenges due to the limited labeled samples and the presence of unknown classes during inference. While recent few-shot learning methods leverage class-level textual information to enhance feature discrimination, they often overlook image-grounded descriptions that provide more fine-grained and instance-specific semantics. Moreover, many of these approaches require prior knowledge of the class name from labeled samples during inference, which may be impractical in open-world scenarios. In this study, we propose a multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance. MGF leverages image descriptions generated by a vision-language model as the textual supervision to guide the learning of semantic features from images. A graph convolutional network is then used to fuse semantic and visual features, enabling effective intra-class information propagation and improving discrimination between known and unknown classes. We jointly optimize a contrastive and a semantic alignment loss to promote intra-class compactness and inter-class separability. Extensive experiments on several few-shot learning benchmarks demonstrate that MGF achieves superior open-set recognition and competitive closed-set classification performance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition. — 科研速览 Science Skim