跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Serli Kopar, Alkis Koudounas, Roshan P. Rane, Sam Gijsen, Paula A. Perez-Toro, Kerstin Ritter·· 15 小时前AI 评分37

超越可解码性:声学因素是否驱动语音阿尔茨海默病评估中的预测?

Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

AI 导读

研究基于 ADReSSo 数据集和三个大型 SSL 骨干模型,对语音施加受控噪声与混响干预,发现声学干预会改变所有三个骨干模型的阿尔茨海默病预测。噪声在原始数据中虽无显著诊断组差异,却产生最强干预效应,且效应方向与分类器决策方向系统性相关,并在留出测试集上复现、在表征空间干预方向反转时随之反转。作者据此主张,干预式鲁棒性测试应成为可信临床语音模型的标准流程。

正文

View PDF HTML (experimental)

Abstract:Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2610.01846 [cs.SD]
  (or arXiv:2610.01846v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.01846

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Serli Kopar [view email]
[v1] Thu, 1 Oct 2026 15:13:19 UTC (156 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org