VGBench 诊断 Audio LLM 的声学上下文门控能力
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
论文提出 VGBench,一个 1,018 项的诊断基准,用于评估 Audio LLM 在旁人说话、自言自语和说话人切换场景下是否应执行动作。六个原始 Audio LLM 和三个免训练适配方法很少在说话人切换时保持静音,最高原始切换静音率仅 14%。
Published on Sep 26
Authors:
,
Abstract
Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each item uses a shared action space comprising silence, a tool call, and a natural-language answer. Speaker-switch pairs hold the specified words fixed while source, distance rendering, and a temporal boundary define a controlled wearer-to-bystander shift. Six raw Audio LLMs and three training-free adaptations often identify the target tool yet rarely withhold action under this shift; the highest raw switch mute rate is 14%. We then use VoxGate as a post-training case study. Supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage has similar switch performance; side-talk accuracy rises from 68.4% to 70.9%, and self-talk muting from 52.0% to 60.0%. Factorized controls identify an independent source-change effect, while sensitivity to the far-field manipulation varies across acoustic renderings. The benchmark therefore measures multi-cue acoustic-context gating rather than isolated speaker identity.
View arXiv page View PDF Add to collection
Get this paper in your agent:
hf papers read 2609.32536
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.32536 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.32536 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.32536 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co