arXiv:cs.AI· Qibing Bai, Yuhan Du, Tom Ko, Shuai Wang, Yannan Wang, Haizhou Li·· 4 小时前AI 评分39
DLM-AN:基于离散扩散的可控口音归一化系统
Controllable Accent Normalization via Discrete Diffusion
AI 导读
DLM-AN 是一个基于掩码离散扩散的可控口音归一化系统,通过 Common Token Predictor 选择性复用源语音 token 来初始化反向扩散过程,复用越多 token 保留的原口音越强。
正文
Abstract:Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control. The implementation is available at this https URL
| Comments: | Accepted to Interspeech 2026 as a long paper |
| Subjects: | Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD) |
| Cite as: | arXiv:2603.14275 [eess.AS] |
| (or arXiv:2603.14275v3 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2603.14275 arXiv-issued DOI via DataCite |
Submission history
From: Qibing Bai [view email]
[v1]
Sun, 15 Mar 2026 08:17:25 UTC (1,634 KB)
[v2]
Mon, 22 Jun 2026 16:56:11 UTC (1,642 KB)
[v3]
Fri, 2 Oct 2026 13:42:06 UTC (1,642 KB)
来源:arXiv:cs.AI · arxiv.org