跳到正文
arXiv:cs.CL· Lukuan Dong, Ziwei Li, Saierdaer Yusuyin, Xianyu Zhao, Zhijian Ou·· 4 小时前AI 评分31

基于 LLM 的多语言音素到字素转换研究:CV-Lang10 基准上 WER 降至 7.66%

Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition

AI 导读

研究者在十语言 CV-Lang10 基准上探索基于 LLM 的多语言音素到字素(P2G)转换,提出 Simplified SKM(S-SKM)蒙特卡洛近似方法,避免 P2G 训练中使用 CTC 的音素概率加权,并引入鲁棒性训练与低资源过采样策略。实验将平均 WER 从 10.56% 降至 7.66%,同时考察了 DANP 等应对 S2P 不确定性的鲁棒性方案。

正文

View PDF HTML (experimental)

Abstract:Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models (LLMs) are promising for P2G, multilingual P2G remains challenging due to language-aware generation and severe cross-language data imbalance. We study multilingual LLM-based P2G on the ten-language CV-Lang10 benchmark. We examine robustness strategies that account for S2P uncertainty, including DANP and Simplified SKM (S-SKM). S-SKM is a Monte Carlo approximation that avoids CTC-based S2P probability weighting in P2G training. Robust training and low-resource oversampling reduce the average WER from 10.56% to 7.66%.
Comments: NCMMSC 2026 camera-ready version
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as: arXiv:2603.29217 [eess.AS]
  (or arXiv:2603.29217v3 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2603.29217

arXiv-issued DOI via DataCite

Submission history

From: Lukuan Dong [view email]
[v1] Tue, 31 Mar 2026 03:32:18 UTC (65 KB)
[v2] Sat, 5 Sep 2026 01:23:42 UTC (62 KB)
[v3] Tue, 6 Oct 2026 18:13:39 UTC (62 KB)

来源:arXiv:cs.CL · arxiv.org