跳到正文
arXiv:cs.AI· Haipeng Liu, Yang Wang, Meng Wang·· 6 小时前AI 评分36

Dual-FDM:在频率感知扩散模型中解耦双重图像参考的个性化生成

Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation

AI 导读

研究者提出 Dual-FDM,一种在频率感知扩散模型中解耦双重图像参考的个性化生成范式,可同时处理定制风格迁移与色彩风格迁移两项任务。该方法通过掩码策略在频域分离不同频段:定制风格迁移用定制参考前景的中频段替换风格参考背景的中频段,色彩风格迁移则用色彩参考前景与背景的低频段替换风格参考背景的低频段,替换频段作为 key 和 value 重建去噪后的查询前景与背景。

正文

View PDF HTML (experimental)

Abstract:Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized this http URL experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from this https URL.
Comments: 28 pages, 14 figures, to appear at NeurIPS 2026, Sydney, Australia
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07684 [cs.CV]
  (or arXiv:2610.07684v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.07684

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haipeng Liu [view email]
[v1] Tue, 6 Oct 2026 03:18:07 UTC (47,767 KB)

来源:arXiv:cs.AI · arxiv.org