arXiv:cs.CL· Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi, Iroro Orife, David Ifeoluwa Adelani·· 4 小时前AI 评分38
ÌròyìnSpeech 文本语料发布:24,905 条约鲁巴语语音与语言技术句料
The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology
AI 导读
ÌròyìnSpeech 发布其文本组件,包含 24,905 条人工核验、标注声调的约鲁巴语句子(275,897 tokens,15,687 types),最初于 2022 年作为录音提示词整理,其 42 小时、80 位说话人的音频自 2024 年起由 ELRA 分发。
正文
Abstract:ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
| Comments: | v2: keywords updated |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.05366 [cs.CL] |
| (or arXiv:2610.05366v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.05366 arXiv-issued DOI via DataCite |
Submission history
From: Kola Tubosun [view email]
[v1]
Sun, 4 Oct 2026 16:47:38 UTC (9 KB)
[v2]
Tue, 6 Oct 2026 19:02:19 UTC (9 KB)
来源:arXiv:cs.CL · arxiv.org