arXiv:cs.CL· Chihiro Taguchi, \'Eric Le Ferrand, Hirosi Nakagawa, Hitomi Ono, Kanji Kato, Emily Prud'hommeaux, David Chiang·· 4 小时前
预训练自监督语音模型 Wav2Vec2 与 HuBERT 能识别未见过的搭嘴音
Pretrained self-supervised speech models can recognize unseen consonants
AI 导读
预训练自监督语音模型 Wav2Vec2 和 HuBERT 在两种搭嘴音丰富的科伊桑语(G|ui 与 West !Xoon)上微调后,识别搭嘴音的准确率始终高于非搭嘴音。研究结果提示,自监督学习能让模型泛化到包括罕见音素在内的人类语音。该工作发表于 Interspeech 2026。
正文
Abstract:Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their training data are heavily skewed toward high-resource languages with little data from low-resource languages, raising concerns about the potential underrepresentation of typologically uncommon speech sounds such as click consonants primarily found in Khoisan languages. This leads to our central research question: Can these models recognize click consonants as accurately as other speech sounds? To address this question, we fine-tune and compare pretrained self-supervised speech models (Wav2Vec2 and HuBERT) on data from two click-rich Khoisan languages (G|ui and West !Xoon). Our results reveal that the fine-tuned models consistently recognize clicks more accurately than non-clicks, suggesting that self-supervision enables generalization across human speech sounds including rare phonemes.
| Comments: | 6 pages, 3 figures, 3 tables, presented at Interspeech 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.11542 [cs.CL] |
| (or arXiv:2606.11542v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.11542 arXiv-issued DOI via DataCite |
Submission history
From: Chihiro Taguchi [view email]
[v1]
Wed, 10 Jun 2026 01:07:32 UTC (162 KB)
[v2]
Thu, 8 Oct 2026 06:36:41 UTC (161 KB)
来源:arXiv:cs.CL · arxiv.org