跳到正文
arXiv:cs.CL· Chihiro Taguchi, \'Eric Le Ferrand, Hirosi Nakagawa, Hitomi Ono, Kanji Kato, Emily Prud'hommeaux, David Chiang·· 4 小时前

预训练自监督语音模型 Wav2Vec2 与 HuBERT 能识别未见过的搭嘴音

Pretrained self-supervised speech models can recognize unseen consonants

AI 导读

预训练自监督语音模型 Wav2Vec2 和 HuBERT 在两种搭嘴音丰富的科伊桑语(G|ui 与 West !Xoon)上微调后,识别搭嘴音的准确率始终高于非搭嘴音。研究结果提示,自监督学习能让模型泛化到包括罕见音素在内的人类语音。该工作发表于 Interspeech 2026。

正文

View PDF HTML (experimental)

Abstract:Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their training data are heavily skewed toward high-resource languages with little data from low-resource languages, raising concerns about the potential underrepresentation of typologically uncommon speech sounds such as click consonants primarily found in Khoisan languages. This leads to our central research question: Can these models recognize click consonants as accurately as other speech sounds? To address this question, we fine-tune and compare pretrained self-supervised speech models (Wav2Vec2 and HuBERT) on data from two click-rich Khoisan languages (G|ui and West !Xoon). Our results reveal that the fine-tuned models consistently recognize clicks more accurately than non-clicks, suggesting that self-supervision enables generalization across human speech sounds including rare phonemes.
Comments: 6 pages, 3 figures, 3 tables, presented at Interspeech 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.11542 [cs.CL]
  (or arXiv:2606.11542v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.11542

arXiv-issued DOI via DataCite

Submission history

From: Chihiro Taguchi [view email]
[v1] Wed, 10 Jun 2026 01:07:32 UTC (162 KB)
[v2] Thu, 8 Oct 2026 06:36:41 UTC (161 KB)

来源:arXiv:cs.CL · arxiv.org