跳到正文
arXiv:cs.CL· Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi, Iroro Orife, David Ifeoluwa Adelani·· 4 小时前AI 评分38

ÌròyìnSpeech 文本语料发布:24,905 条约鲁巴语语音与语言技术句料

The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology

AI 导读

ÌròyìnSpeech 发布其文本组件,包含 24,905 条人工核验、标注声调的约鲁巴语句子(275,897 tokens,15,687 types),最初于 2022 年作为录音提示词整理,其 42 小时、80 位说话人的音频自 2024 年起由 ELRA 分发。

正文

View PDF HTML (experimental)

Abstract:ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
Comments: v2: keywords updated
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.05366 [cs.CL]
  (or arXiv:2610.05366v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.05366

arXiv-issued DOI via DataCite

Submission history

From: Kola Tubosun [view email]
[v1] Sun, 4 Oct 2026 16:47:38 UTC (9 KB)
[v2] Tue, 6 Oct 2026 19:02:19 UTC (9 KB)

来源:arXiv:cs.CL · arxiv.org