arXiv:cs.CL· Shunsuke Mitsumori, Tomoya Mizumoto, Yusuke Fujita·· 3 小时前AI 评分34
用合成伪方言增强数据提升语音语言模型的方言鲁棒性
Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation
AI 导读
研究者提出用 LLM 生成方言文本、再经标准语 TTS 模型合成伪方言语音,无需真实方言语音即可增强语音语言模型(SLM)的方言理解能力。在日语、德语、汉语方言到英语的语音翻译评测中,伪方言增强使日语得分从 25.38 升至 26.24、德语从 31.57 升至 32.47;加入中间标准文本预测后,日语进一步升至 28.26、汉语从 11.67 升至 16.37。
正文
Abstract:Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.
| Comments: | 7 pages, 1 figure, 7 tables. Accepted to IEEE SLT 2026 |
| Subjects: | Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.09321 [cs.CL] |
| (or arXiv:2610.09321v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09321 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shunsuke Mitsumori [view email]
[v1]
Wed, 7 Oct 2026 02:25:43 UTC (190 KB)
来源:arXiv:cs.CL · arxiv.org