跳到正文
arXiv:cs.CL· Seymanur Akti, Alexander Waibel·· 6 小时前AI 评分35

F5-TTS 零样本 Lombard 语音合成:用可控风格嵌入调控发声力度

Zero-Shot Lombard Speech Synthesis with Controllable Style Embeddings

AI 导读

研究者扩展 F5-TTS,提出一种零样本合成 Lombard 语音的可控 TTS 系统,无需 Lombard 专用训练数据。该方法引入可学习的风格嵌入,并用 PCA 分析其潜空间以定位与 Lombard 属性相关的方向,从而对发声力度和发音清晰度实现可解释控制并生成不同 Lombard 等级语音。实验显示该方法保持说话人身份与自然度,提升噪声下的可懂度,并可泛化到未见过的说话人。

正文

View PDF HTML (experimental)

Abstract:The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard-like speech in a zero-shot manner without requiring Lombard-specific training data. Our approach extends F5-TTS with a learned style embedding representation and analyzes the resulting latent space using principal component analysis (PCA) to identify directions associated with Lombard-related attributes. By manipulating these directions, we obtain interpretable control over vocal effort and articulation and generate speech at different Lombard levels. Experimental results show that the proposed method preserves speaker identity and naturalness, improves intelligibility under noisy conditions, and generalizes to previously unseen speakers. These findings demonstrate that style-embedding manipulation provides an effective and scalable framework for controllable zero-shot Lombard speech synthesis.
Comments: Accepted at IEEE SLT 2026
Subjects: Sound (cs.SD); Computation and Language (cs.CL)
Cite as: arXiv:2601.12966 [cs.SD]
  (or arXiv:2601.12966v2 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2601.12966

arXiv-issued DOI via DataCite

Submission history

From: Şeymanur Aktı [view email]
[v1] Mon, 19 Jan 2026 11:25:19 UTC (6,347 KB)
[v2] Tue, 6 Oct 2026 02:32:16 UTC (6,013 KB)

来源:arXiv:cs.CL · arxiv.org