arXiv:cs.AI· David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus·· 5 小时前AI 评分39
基于人工对话协议的构音障碍与气管造口说话人个性化 ASR 系统
Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations
AI 导读
研究团队为一名患有永久性气管造口和严重构音障碍的捷克语说话人构建了个性化 ASR 系统,并公开了 33 小时标注语音数据集,该数据集通过新型“人工对话”协议采集。
正文
Abstract:This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.
| Comments: | 8 pages, three figures, to be published in IEEE Speech Language Technology workshop 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Sound (cs.SD) |
| Cite as: | arXiv:2610.03017 [cs.AI] |
| (or arXiv:2610.03017v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03017 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: David Nadrchal [view email]
[v1]
Fri, 2 Oct 2026 08:52:00 UTC (910 KB)
来源:arXiv:cs.AI · arxiv.org