arXiv:cs.CL· Seymanur Akti, Alexander Waibel·· 6 小时前AI 评分39
Loud and Clear:用动态激活引导让 TTS 模型在噪声环境中说出更清晰的语音
Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments
AI 导读
研究提出一种 prompt-relative 激活引导机制,无需重新训练即可动态控制预训练 TTS 模型,模拟 Lombard 效应中的发声用力增强与超清晰发音。在已见和未见说话人及多语言测试中,该方法系统改变 Lombard 相关声学特征,保持 89-95% 的说话人相似度,并在 1 dB SNR 背景噪声下将 WER 降低 7-22%。
正文
Abstract:Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.
| Comments: | Submitted to ICASSP 2027 |
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.07647 [cs.SD] |
| (or arXiv:2610.07647v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07647 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Şeymanur Aktı [view email]
[v1]
Tue, 6 Oct 2026 02:41:15 UTC (832 KB)
来源:arXiv:cs.CL · arxiv.org