科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-09-02

Optimal Transport-based Difficulty-aware Contribution Allocation for Dataset Distillation.

Xiao Cui, Yulei Qin, Mo Zhu, Wengang Zhou, Yu Zhu, Hongsheng Li, Ming-Hsuan Yang, Houqiang Li

原始摘要(英文原文)· Original abstract
The increasing reliance on large-scale datasets imposes significant storage and computational burdens on training deep learning models. Dataset distillation methods, particularly those based on sample generation, aim to condense large original datasets into compact synthetic sets while preserving essential information. Existing subset synthesis approaches typically minimize a homogeneous distance, assigning uniform contributions from all real instances to the construction of each synthetic sample. We show that such equal allocation neglects instance-level relationships between real-synthetic pairs, leading to inadequate modeling of the geometric structural discrepancies between the distilled and original datasets. In this work, we reformulate homogeneous distance minimization as a bi-level optimization problem via a matching-and-approximating paradigm. In the matching stage, we employ an optimal transport matrix to dynamically allocate contributions from real instances. Building upon this transport-based allocation, we further introduce a difficulty-aware marginal reweighting mechanism to emphasize informative instances while preserving global geometric consistency. In the subsequent approximation stage, synthetic samples are refined according to the established allocation scheme to better approximate the real data distribution. This strategy enables a more faithful characterization of intricate geometric structures and improved handling of intra-class variations, thereby enhancing distillation fidelity. Extensive experiments across diverse architectures, modalities, and learning paradigms demonstrate that the proposed framework consistently improves performance, with gains observed in standard supervised, federated, continual, and multimodal learning settings.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Optimal Transport-based Difficulty-aware Contribution Allocation for Dataset Distillation. — 科研速览 Science Skim