跳到正文
arXiv:cs.CL· Ryo Magoshi, Shinsuke Sakai, Tatsuya Kawahara·· 3 小时前AI 评分34

面向 LLM 语音识别的音素引导初始化方法

Phoneme-Guided Initialization for LLM-based Speech Recognition

AI 导读

研究者提出音素引导初始化:先在语音转音素(S2P)任务上预训练音频编码器、在音素转字形(P2G)任务上预训练 LLM,再拼接端到端微调目标 ASR 任务。在日语(CSJ)、中文(AISHELL-1)及 Common Voice 25.0 的鞑靼语、乌尔都语两项低资源语言上,该方法持平或优于级联 S2P-P2G 基线与未做 P2G 初始化的端到端模型,论文已被 IEEE SLT 2026 接收。

正文

View PDF HTML (experimental)

Abstract:Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.
Comments: Accepted at IEEE SLT 2026
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as: arXiv:2610.08994 [eess.AS]
  (or arXiv:2610.08994v1 [eess.AS] for this version)
  https://doi.org/10.48550/arXiv.2610.08994

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ryo Magoshi [view email]
[v1] Tue, 6 Oct 2026 18:52:05 UTC (394 KB)

来源:arXiv:cs.CL · arxiv.org