跳到正文
arXiv:cs.LG· Kanthila Chinmayi (IRDL, LaTIM), Abdallah Nassib (LARIS), Misy Harrison (LaTIM), Panheleux Celine (LaTIM, CHU - BREST), Saliou Vanessa (CHU - BREST), Seizeur Romuald (LaTIM), Dardenne Guillaume (LaTIM)·· 3 小时前

自监督语音表征用于清醒开颅术中跨说话人构音障碍检测

Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy

AI 导读

一项针对清醒开颅术中构音障碍检测的系统性组件评估显示,语音表征对跨说话人性能的影响远大于分类器选择:三种分类器 AUC 差距不超过 4.7%,而用多层 wav2vec 2.0 自监督表征替换传统声学特征后 AUC 提升 18.2%-26.1%。

正文

View PDF

Abstract:Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function. However, automated detection remains challenging because operating-room recordings contain substantial acoustic interference, clinically relevant speech events are rare, and available cohorts are small and heterogeneous across speakers. This study presents a systematic component-wise evaluation of a pipeline for distinguishing dysarthric from no-trouble speech in the DATABRASE corpus of awake-craniotomy recordings. The pipeline incorporates speaker diarization to isolate patient speech, a multi-view representation combining handcrafted acoustic descriptors with multilayer wav2vec 2.0 embeddings, speaker-conditional normalization and transferability-based feature selection to improve cross-speaker robustness, and a cascaded classifier comprising a gradient-boosted first stage and a neural second stage. Evaluation was conducted under strict speaker-independent conditions using leave-one-speaker-out cross-validation. The results show that cross-speaker performance is influenced more strongly by the speech representation than by classifier choice. The AUCs of three classifiers differed by no more than 4.7%, whereas replacing conventional acoustic descriptors with the multilayer self-supervised representation produced AUC improvements of 18.2%-26.1%. Diarization-conditioned feature extraction and the proposed classifier cascade provided additional consistent gains. These findings indicate that reliable patient-specific speech isolation and strong pretrained representations are more important than increased classifier complexity in low-resource intra-operative settings. They also quantify the potential performance gains that may be achieved through patient-specific preoperative calibration.
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
Cite as: arXiv:2610.11825 [cs.LG]
  (or arXiv:2610.11825v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11825

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: KANTHILA Chinmayi [view email] [via CCSD proxy]
[v1] Thu, 8 Oct 2026 12:25:15 UTC (2,221 KB)

来源:arXiv:cs.LG · arxiv.org