arXiv:cs.CL· Mengzhe Geng·· 4 小时前AI 评分37
生成式音频调用对已知任务音频 LLM 评测的增量价值审计
Auditing generative audio calls for known-task audio-llm evaluation
AI 导读
一项审计研究用匹配选择器估计生成式音频调用的增量准确率:在 VocalSound 上,仅转写准确率为 0.296,而受监督的 CLAP 和 WavLM 控制组无需调用即达到 0.850 和 0.854。
正文
Abstract:Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call selector (difference 0.004; 95% CI [-0.025, 0.033]). Thus, results do not show a call gain after transcript and encoder evidence are available. Relevant quantity is incremental accuracy from allowing calls, not the waveform-transcript gap.
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2608.27817 [cs.SD] |
| (or arXiv:2608.27817v4 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2608.27817 arXiv-issued DOI via DataCite |
Submission history
From: Mengzhe Geng [view email]
[v1]
Fri, 28 Aug 2026 01:33:47 UTC (18 KB)
[v2]
Mon, 31 Aug 2026 19:27:04 UTC (18 KB)
[v3]
Sat, 12 Sep 2026 22:36:51 UTC (36 KB)
[v4]
Wed, 7 Oct 2026 11:58:34 UTC (49 KB)
来源:arXiv:cs.CL · arxiv.org