arXiv:cs.LG· Kanthila Chinmayi (IRDL, LaTIM), Abdallah Nassib (LARIS), Misy Harrison (LaTIM), Panheleux Celine (LaTIM, CHU - BREST), Saliou Vanessa (CHU - BREST), Seizeur Romuald (LaTIM), Dardenne Guillaume (LaTIM)·· 3 小时前
自监督语音表征用于清醒开颅术中跨说话人构音障碍检测
Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy
AI 导读
一项针对清醒开颅术中构音障碍检测的系统性组件评估显示,语音表征对跨说话人性能的影响远大于分类器选择:三种分类器 AUC 差距不超过 4.7%,而用多层 wav2vec 2.0 自监督表征替换传统声学特征后 AUC 提升 18.2%-26.1%。
正文
Abstract:Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function. However, automated detection remains challenging because operating-room recordings contain substantial acoustic interference, clinically relevant speech events are rare, and available cohorts are small and heterogeneous across speakers. This study presents a systematic component-wise evaluation of a pipeline for distinguishing dysarthric from no-trouble speech in the DATABRASE corpus of awake-craniotomy recordings. The pipeline incorporates speaker diarization to isolate patient speech, a multi-view representation combining handcrafted acoustic descriptors with multilayer wav2vec 2.0 embeddings, speaker-conditional normalization and transferability-based feature selection to improve cross-speaker robustness, and a cascaded classifier comprising a gradient-boosted first stage and a neural second stage. Evaluation was conducted under strict speaker-independent conditions using leave-one-speaker-out cross-validation. The results show that cross-speaker performance is influenced more strongly by the speech representation than by classifier choice. The AUCs of three classifiers differed by no more than 4.7%, whereas replacing conventional acoustic descriptors with the multilayer self-supervised representation produced AUC improvements of 18.2%-26.1%. Diarization-conditioned feature extraction and the proposed classifier cascade provided additional consistent gains. These findings indicate that reliable patient-specific speech isolation and strong pretrained representations are more important than increased classifier complexity in low-resource intra-operative settings. They also quantify the potential performance gains that may be achieved through patient-specific preoperative calibration.
| Subjects: | Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP) |
| Cite as: | arXiv:2610.11825 [cs.LG] |
| (or arXiv:2610.11825v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11825 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: KANTHILA Chinmayi [view email] [via CCSD proxy]
[v1]
Thu, 8 Oct 2026 12:25:15 UTC (2,221 KB)
来源:arXiv:cs.LG · arxiv.org