Model Distribution-Aware Multimodal Dataset Distillation
Jiajun Shen ⋅ Yonglin Wu ⋅ Ruonan Yu ⋅ Xinchao Wang
Abstract
Multimodal dataset distillation (MDD) seeks to synthesize compact surrogates from large-scale image-text corpora, reducing the high cost of pre-training. Although recent feature-matching methods have improved the efficiency of MDD, they often suffer from initialization overfitting, which distorts the cross-modal structure and severely degrades retrieval performance under unseen initializations. To mitigate this issue, we propose $\textbf{MDA-DD}$, which enhances generalization via a dynamic staggered model queue. We further identify two architectural bottlenecks in prior work: structurally redundant projections and inefficient covariance computation. For the former, we design an asymmetric dual projection to reduce redundancy while preserving balanced bidirectional cross-modal alignment. For the latter, we introduce a Second-Order Cross-Modal Moment (SOCM) objective, which efficiently captures both mean and correlation statistics across modalities. Experiments on Flickr30K and MS-COCO show that $\textbf{MDA-DD}$ consistently surpasses existing methods, yielding notable retrieval improvements (e.g., +12.9 IR@1 and +18.2 TR@1 in the 100-pair setting) and matching or exceeding baselines trained on 500 pairs using only 100 pairs.
Chat is not available.
Successful Page Load