跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, Ren\'{e} S Kahn, Guillermo Checci, Baihan Lin·· 5 小时前AI 评分31

用音频语言模型验证精神病学语音记录中的角色引导说话人删除

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

AI 导读

研究用音频语言模型验证精神病学双人访谈录音中指定说话人(临床医生或患者)的语音是否已被删除,在48段录音上测试了Gemma-4-12B、Gemma-4-31B、Nemotron-3-Nano和Nemotron-3-Nano-Omni四个开放权重模型。十四种模型-视图配置的OR集成取得F1 0.478(精确率0.330、召回率0.870),召回率提升显示模型与上下文视图间存在显著互补性。

正文

View PDF HTML (experimental)

Abstract:Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.
Subjects: Machine Learning (cs.LG); Sound (cs.SD)
Cite as: arXiv:2609.38491 [cs.LG]
  (or arXiv:2609.38491v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38491

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Joseph Colonel [view email]
[v1] Tue, 29 Sep 2026 20:20:24 UTC (24 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org