跳到正文
arXiv:cs.AI· Wenjie Zheng, Haoji Hu, Jiali Lu, Xingze Zou, Jing Wang·· 6 小时前AI 评分36

D3S2:面向语义分割的扩散引导数据集蒸馏框架

D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation

AI 导读

研究者提出 D3S2,一个面向语义分割的扩散引导数据集蒸馏框架,通过类平衡掩码选择与扩散引导图像合成两阶段设计,解决长尾类不平衡与像素级对齐问题。在 1% 压缩率下,D3S2 配合 Mask2Former (Swin-S) 在 ADE20K 和 COCO-Stuff 上分别达到 24.99% 和 35.49% mIoU,比随机选择高出 9.34% 和 5.70%。代码已开源。

正文

View PDF HTML (experimental)

Abstract:Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (i) long-tailed class imbalance, (ii) the need for strict pixel-wise alignment between images and dense labels, and (iii) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose D3S2, a Diffusion-guided Dataset Distillation framework for Semantic Segmentation. Our method adopts a two-stage design. In Class-Balanced Mask Selection, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In Diffusion-Guided Image Synthesis, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of D3S2. Notably, at an extremely compression rate of 1%, our method achieves 24.99% and 35.49% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34% and 5.70%, respectively. Our code is available at this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.25022 [cs.CV]
  (or arXiv:2605.25022v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2605.25022

arXiv-issued DOI via DataCite

Submission history

From: Wenjie Zheng [view email]
[v1] Sun, 24 May 2026 12:01:38 UTC (21,256 KB)
[v2] Tue, 6 Oct 2026 09:48:58 UTC (22,633 KB)

来源:arXiv:cs.AI · arxiv.org