arXiv:cs.CL· Lukuan Dong, Ziwei Li, Saierdaer Yusuyin, Xianyu Zhao, Zhijian Ou·· 4 小时前AI 评分31
基于 LLM 的多语言音素到字素转换研究:CV-Lang10 基准上 WER 降至 7.66%
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
AI 导读
研究者在十语言 CV-Lang10 基准上探索基于 LLM 的多语言音素到字素(P2G)转换,提出 Simplified SKM(S-SKM)蒙特卡洛近似方法,避免 P2G 训练中使用 CTC 的音素概率加权,并引入鲁棒性训练与低资源过采样策略。实验将平均 WER 从 10.56% 降至 7.66%,同时考察了 DANP 等应对 S2P 不确定性的鲁棒性方案。
正文
Abstract:Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
| Comments: | NCMMSC 2026 camera-ready version |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD) |
| Cite as: | arXiv:2603.29217 [eess.AS] |
| (or arXiv:2603.29217v3 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2603.29217 arXiv-issued DOI via DataCite |
Submission history
From: Lukuan Dong [view email]
[v1]
Tue, 31 Mar 2026 03:32:18 UTC (65 KB)
[v2]
Sat, 5 Sep 2026 01:23:42 UTC (62 KB)
[v3]
Tue, 6 Oct 2026 18:13:39 UTC (62 KB)
来源:arXiv:cs.CL · arxiv.org