跳到正文
arXiv:cs.CL· Hieu Hoang, Amittai Axelrod·· 4 小时前AI 评分33

用部分语音学习何时提交:面向端到端同声传译的 prefix 监督

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

AI 导读

研究者用语音语言模型自身完整与部分波形的翻译结果构造 prefix 监督,无需转录文本或人工翻译,即可适配端到端同声传译。

正文

View PDF HTML (experimental)

Abstract:Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.02612 [cs.CL]
  (or arXiv:2610.02612v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02612

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hieu Hoang [view email]
[v1] Fri, 2 Oct 2026 00:13:52 UTC (713 KB)

来源:arXiv:cs.CL · arxiv.org