arXiv:cs.LG· Marcel Gibier, Thomas Thebaud, Olivier Bo\"effard, Jean-Fran\c{c}ois Bonastre·· 4 小时前AI 评分33
CARES:面向说话者对声音反应的可控合成基准
CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound
AI 导读
研究团队构建了 CARES 基准,包含 10,000 个双说话人场景,用"说话者能听出反应"这一规则定义声音显著性,并由语言模型生成对话以固定标注答案。团队随后在场景识别、声音打标和反应分类三项任务上评测了六个音频语言模型,结果显示这些模型能听到声音,却识别不出说话者对其的反应。
正文
Abstract:Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
| Comments: | Submitted to ICASSP 2027 |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10208 [cs.SD] |
| (or arXiv:2610.10208v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10208 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marcel Gibier [view email]
[v1]
Wed, 7 Oct 2026 15:05:59 UTC (172 KB)
来源:arXiv:cs.LG · arxiv.org