arXiv:cs.CL· Ryo Magoshi, Shinsuke Sakai, Tatsuya Kawahara·· 3 小时前AI 评分34
面向 LLM 语音识别的音素引导初始化方法
Phoneme-Guided Initialization for LLM-based Speech Recognition
AI 导读
研究者提出音素引导初始化:先在语音转音素(S2P)任务上预训练音频编码器、在音素转字形(P2G)任务上预训练 LLM,再拼接端到端微调目标 ASR 任务。在日语(CSJ)、中文(AISHELL-1)及 Common Voice 25.0 的鞑靼语、乌尔都语两项低资源语言上,该方法持平或优于级联 S2P-P2G 基线与未做 P2G 初始化的端到端模型,论文已被 IEEE SLT 2026 接收。
正文
Abstract:Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.
| Comments: | Accepted at IEEE SLT 2026 |
| Subjects: | Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD) |
| Cite as: | arXiv:2610.08994 [eess.AS] |
| (or arXiv:2610.08994v1 [eess.AS] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08994 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ryo Magoshi [view email]
[v1]
Tue, 6 Oct 2026 18:52:05 UTC (394 KB)
来源:arXiv:cs.CL · arxiv.org