科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-08-31

Pose-Star++: Semantic-Visual Understanding for Fine-Grained Fashion Image Editing.

Yuran Dong, Bo Du, Mang Ye

原始摘要(英文原文)· Original abstract
Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are semantically flexible but structurally inaccurate. To bridge this gap, we propose Pose-Star++, a training-free, plug-and-play framework that introduces two core innovations: an LVLM-based Understanding Module that shifts from word- to sentence-level semantic-visual comprehension, eliminating cumbersome instruction pre-parsing and enabling robust understanding of complex natural language; a Bidirectional Calibration Module that co-optimizes semantic and structural constraints through forward pose-guided and backward attention-guided refinement, achieving precise, whole-body-reachable region calibration even under challenging in-the-wild poses. We further contribute the first real-world-oriented fashion-editing benchmark with diverse data, instructions, and tasks, exposing long-overlooked practical challenges. Extensive experiments demonstrate that Pose-Star++ significantly outperforms existing methods in semantic alignment, pose robustness, and in-the-wild generalization across complex scenarios, advancing toward practical, user-guided fashion creation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Pose-Star++: Semantic-Visual Understanding for Fine-Grained Fashion Image Editing. — 科研速览 Science Skim