跳到正文
arXiv:cs.LG· Marcel Gibier, Thomas Thebaud, Olivier Bo\"effard, Jean-Fran\c{c}ois Bonastre·· 4 小时前AI 评分33

CARES:面向说话者对声音反应的可控合成基准

CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound

AI 导读

研究团队构建了 CARES 基准,包含 10,000 个双说话人场景,用"说话者能听出反应"这一规则定义声音显著性,并由语言模型生成对话以固定标注答案。团队随后在场景识别、声音打标和反应分类三项任务上评测了六个音频语言模型,结果显示这些模型能听到声音,却识别不出说话者对其的反应。

正文

View PDF HTML (experimental)

Abstract:Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
Comments: Submitted to ICASSP 2027
Subjects: Sound (cs.SD); Machine Learning (cs.LG)
Cite as: arXiv:2610.10208 [cs.SD]
  (or arXiv:2610.10208v1 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2610.10208

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Marcel Gibier [view email]
[v1] Wed, 7 Oct 2026 15:05:59 UTC (172 KB)

来源:arXiv:cs.LG · arxiv.org