跳到正文
arXiv:cs.CL· Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson, Kishankumar Bhimani, Jochebed Afua Basil·· 4 小时前AI 评分36

MGhana-ST:面向加纳语言的低资源语音翻译数据集及多语言训练权衡分析

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

AI 导读

研究者发布 MGhana-ST 语音翻译数据集,覆盖 Ga、Twi(Akuapem 与 Asante)、Ewe、Fante 四种加纳低资源语言变体,实验使用约 16.1 小时语音与英语译文配对数据,译文由 37 名母语标注者直接从音频产出。

正文

View PDF HTML (experimental)

Abstract:We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The audio is curated from two existing Ghanaian speech resources. Unlike in those resources, the English translations are produced directly from audio by 37 native-speaker annotators and include verbal and non-verbal event annotations.
Using Whisper-small, we compare monolingual and multilingual training under severe data scarcity, reporting means over three seeds. Flat multilingual training benefits no variety in this regime. Ga and Twi are unchanged within seed variance (+0.51 and +0.06 BLEU against monolingual standard deviations of 1.63 and 2.20), while Ewe declines by 6.99 BLEU and Fante by 5.11. The degrading varieties are Ewe, which is linguistically distinct and drawn from a different source corpus, and Fante, the least-resourced. Comparing empirical cross-lingual transfer with typology-based similarity, we find that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training.
We also report a methodological finding. An earlier single-run analysis found positive transfer for three of four varieties; this did not survive replication across seeds. For Ga and Twi, monolingual baselines trained on 1.6 to 6.2 hours of audio have seed standard deviations roughly five and thirty times those of the multilingual models (0.35 and 0.07 BLEU). When the monolingual condition is noisier, a single-run comparison can show apparent transfer of this size from seed variation alone. We release MGhana-ST to support research on African language speech technology and low-resource speech translation.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.40041 [cs.CL]
  (or arXiv:2609.40041v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.40041

arXiv-issued DOI via DataCite

Submission history

From: Frank Lawrence Nii Adoquaye Acquaye [view email]
[v1] Wed, 30 Sep 2026 16:10:35 UTC (113 KB)
[v2] Thu, 1 Oct 2026 21:48:09 UTC (113 KB)

来源:arXiv:cs.CL · arxiv.org