arXiv:cs.CL· Moran Yanuka, Raja Giryes, Moris Alper·· 3 小时前
多语言语音模型中的音系干扰问题
Phonological Interference in Multilingual Speech Models
AI 导读
音素级语音模型存在一种系统性失败模式——音系干扰:模型默认输入为单一语言并强加其音系,覆盖与之冲突的音素决策。在语码切换输入上,两个音素识别器和一个音素条件 TTS 模型丢失了 32% 至 79% 的此类音素,而对两种语言共有的音素丢失远少。研究者提出窗口化语言估计(WLE),在推理时修复语言估计,在三种模型上消除 34% 至 69% 的干扰,且识别器的单语性能基本不变。
正文
Abstract:Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.
| Comments: | Preprint |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11275 [cs.CL] |
| (or arXiv:2610.11275v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11275 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Moran Yanuka [view email]
[v1]
Thu, 8 Oct 2026 05:38:01 UTC (272 KB)
来源:arXiv:cs.CL · arxiv.org