CoRe-Diffusion: Bridging the Generative-Discriminative Gap in Dataset Distillation via Manifold-Aligned Contrastive Guidance

Authors: Heng Shu, Linjuan Cheng, Qiming Yang, Yuquan Wu
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Generative Dataset Distillation, Contrastive Negative Guidance, Manifold Alignment, Diffusion Models

Abstract

Dataset Distillation aims to condense large-scale datasets into a tiny but highly informative subset, enabling efficient training while achieving performance comparable to the full dataset. To overcome the scalability bottlenecks of traditional optimization, Generative Dataset Distillation (GDD) leverages diffusion priors to reformulate the dataset condensation process as an efficient generative sampling paradigm. However, existing GDD paradigms primarily optimize intra-class likelihood while largely overlooking explicit inter-class separation, which may cause generated features to concentrate near decision boundaries. Furthermore, relying on manually designed guidance targets to steer synthesis often pushes the diffusion trajectory off the natural data manifold, inducing severe structural artifacts. To address these issues, CoRe-Diffusion is proposed as a unified framework. Specifically, it introduces Contrastive Negative Guidance, which utilizes inter-class mode centers—computed but largely unused during synthesis by existing methods—to exert zero-overhead repulsive gradients against hard negatives. To preserve synthesis fidelity, Manifold-Aligned Real-Anchor Discovery strictly constrains the guidance targets to the exact latent representations of real images, preventing the trajectory from drifting. This intrinsically preserves complex semantic structures, bypassing the prohibitive computational overhead of relying on external generative priors or auxiliary modules. Alongside an Annealed Sampling Schedule for smooth trajectory evolution, CoRe-Diffusion achieves state-of-the-art performance, yielding a 3.8\% absolute accuracy improvement on ImageNet-1K while introducing negligible computational overhead.
📄 View Full Paper (PDF) 📋 Show Citation