arXiv:cs.CL· Long-Vu Hoang, Naomi Harte·· 4 小时前AI 评分38
鸡尾酒会场景下的音视频轮次转换预测研究
Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
AI 导读
研究评估了在干净数据上训练的音视频轮次转换预测模型(PTTM)在基于 AVCocktail 数据集构建的鸡尾酒会测试平台上的表现,发现噪声条件下音频和视觉模态性能一致下降,加权 F1 最多相对降低 38%。在新领域上微调可提升鲁棒性,但增益因模态而异,且取决于可用预训练数据的规模。所有代码和轮次标注已公开。
正文
Abstract:Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.
| Comments: | Accepted to IEEE SLT 2026. This version includes an appendix about manual verified labels for AVCocktail |
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.17056 [cs.SD] |
| (or arXiv:2609.17056v2 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2609.17056 arXiv-issued DOI via DataCite |
Submission history
From: Vu Hoang [view email]
[v1]
Tue, 15 Sep 2026 12:06:16 UTC (9,287 KB)
[v2]
Thu, 1 Oct 2026 22:22:55 UTC (9,287 KB)
来源:arXiv:cs.CL · arxiv.org