arXiv:cs.LG· Jingqi Sun, Haozhan Tang, Shulin He, Zhong-Qiu Wang·· 4 小时前AI 评分27
VM-ArrayDPS:虚拟麦克风增强扩散后验采样用于无监督盲语音分离
VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation
AI 导读
VM-ArrayDPS 通过为麦克风阵列增补高信噪比的虚拟麦克风,为 ArrayDPS 的后验采样提供额外多通道一致性约束,从而提升无监督盲语音分离性能。在 2 说话人和 3 说话人数据集上,该方法显著优于 ArrayDPS,消融实验进一步考察了虚拟麦克风数量及虚拟麦克风所带来多通道一致性目标权重的影响。该工作已投稿 ICASSP 2027,目前处于审稿阶段。
正文
Abstract:Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.
| Comments: | Submitted to ICASSP 2027 and currently under review |
| Subjects: | Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Signal Processing (eess.SP) |
| Cite as: | arXiv:2610.09334 [eess.AS] |
| (or arXiv:2610.09334v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09334 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jingqi Sun [view email]
[v1]
Wed, 7 Oct 2026 02:50:56 UTC (90 KB)
来源:arXiv:cs.LG · arxiv.org