跳到正文
arXiv:cs.LG· Xiao Zhang, Yuxin Chen, Zhixuan Liang, Guojian Zhan, Chenran Li, Chenfeng Xu, Masayoshi Tomizuka, Yiheng Li·· 4 小时前AI 评分44

Immiscible Diffusion Policy:无需标签的噪声分配如何保住机器人多模态动作

Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment

AI 导读

研究者提出 Immiscible Diffusion Policy,一种无需标签的训练期插件,通过动作-噪声分配保留相对独立的噪声到动作路径,不改变策略架构与推理流程,以缓解扩散策略坍缩到单一模态的问题。在五个仿真和两个真实世界人形操作任务上,该方法将三个双模态任务中非主导模态占比提升 6.0x-14.6x,并在两个四模态任务中找回原策略 rollout 中完全缺失的演示模态。

正文

View PDF HTML (experimental)

Abstract:When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.
Subjects: Robotics (cs.RO); Machine Learning (cs.LG)
Cite as: arXiv:2610.09369 [cs.RO]
  (or arXiv:2610.09369v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2610.09369

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xiao Zhang [view email]
[v1] Wed, 7 Oct 2026 03:22:26 UTC (2,865 KB)

来源:arXiv:cs.LG · arxiv.org