跳到正文
arXiv:cs.CL· Mengzhe Geng·· 4 小时前AI 评分37

生成式音频调用对已知任务音频 LLM 评测的增量价值审计

Auditing generative audio calls for known-task audio-llm evaluation

AI 导读

一项审计研究用匹配选择器估计生成式音频调用的增量准确率:在 VocalSound 上,仅转写准确率为 0.296,而受监督的 CLAP 和 WavLM 控制组无需调用即达到 0.850 和 0.854。

正文

View PDF HTML (experimental)

Abstract:Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call selector (difference 0.004; 95% CI [-0.025, 0.033]). Thus, results do not show a call gain after transcript and encoder evidence are available. Relevant quantity is incremental accuracy from allowing calls, not the waveform-transcript gap.
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
Cite as: arXiv:2608.27817 [cs.SD]
  (or arXiv:2608.27817v4 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2608.27817

arXiv-issued DOI via DataCite

Submission history

From: Mengzhe Geng [view email]
[v1] Fri, 28 Aug 2026 01:33:47 UTC (18 KB)
[v2] Mon, 31 Aug 2026 19:27:04 UTC (18 KB)
[v3] Sat, 12 Sep 2026 22:36:51 UTC (36 KB)
[v4] Wed, 7 Oct 2026 11:58:34 UTC (49 KB)

来源:arXiv:cs.CL · arxiv.org