跳到正文
arXiv:cs.CL· Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh, Bhiksha Raj·· 4 小时前AI 评分34

CoLMbo-SV:面向可解释说话人验证的接地语言模型

CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification

AI 导读

CoLMbo-SV 将预训练说话人编码器接入语言模型并输入显式声学测量,在 VoxCeleb1-O 上实现 0.99% EER,较在 VoxReason 上微调的最强音频语言基线降低约 80% 验证错误,数值接地得分 0.82。研究同时发布 VoxReason 配对录音数据集,提供经数值与定性校验的比较报告监督。分析显示声学正确性与决策相关性是解释的两个独立属性,数值接地指标未能覆盖这一差距。

正文

View PDF HTML (experimental)

Abstract:Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2609.33212 [cs.CL]
  (or arXiv:2609.33212v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.33212

arXiv-issued DOI via DataCite

Submission history

From: Massa Baali [view email]
[v1] Sun, 27 Sep 2026 04:51:53 UTC (5,033 KB)
[v2] Fri, 2 Oct 2026 16:04:06 UTC (5,033 KB)

来源:arXiv:cs.CL · arxiv.org