arXiv:cs.LG· Xiao Zhang, Yuxin Chen, Zhixuan Liang, Guojian Zhan, Chenran Li, Chenfeng Xu, Masayoshi Tomizuka, Yiheng Li·· 4 小时前AI 评分44
Immiscible Diffusion Policy:无需标签的噪声分配如何保住机器人多模态动作
Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment
AI 导读
研究者提出 Immiscible Diffusion Policy,一种无需标签的训练期插件,通过动作-噪声分配保留相对独立的噪声到动作路径,不改变策略架构与推理流程,以缓解扩散策略坍缩到单一模态的问题。在五个仿真和两个真实世界人形操作任务上,该方法将三个双模态任务中非主导模态占比提升 6.0x-14.6x,并在两个四模态任务中找回原策略 rollout 中完全缺失的演示模态。
正文
Abstract:When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.
| Subjects: | Robotics (cs.RO); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09369 [cs.RO] |
| (or arXiv:2610.09369v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09369 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiao Zhang [view email]
[v1]
Wed, 7 Oct 2026 03:22:26 UTC (2,865 KB)
来源:arXiv:cs.LG · arxiv.org