Jianhui Kong, Sitong Ye, Hangming Xie, Yuying He
Clothing image retrieval methods often emphasize color and texture, which can lead to confusion between garments that share a similar appearance but differ in their component layouts. This study presents a component-aware feature-learning protocol for garment image retrieval and design reference comparison. The protocol parses semantic garment regions, constructs an image-level garment component graph, encodes appearance and component relations in parallel, and fuses both representations using structural consistency, structural relationship, and structural discrimination losses. DeepFashion2 images are standardized to 512 x 512 pixels, parsed with HRNet-W48, encoded with ResNet-50 and graph convolution, and evaluated through structural retrieval and clustering. The complete model achieved a structural consistency index of 0.81, a structural discrimination index of 0.71, a Recall@5 of 89.2%, a mean average precision of 86.4%, and a silhouette coefficient of 0.74. The protocol supports garment structure retrieval, structural comparison, and difference localization when surface appearance does not reliably represent garment construction.