arXiv:cs.AI· Yu Chen, Ruihang Liu, Yangguang Xu, Xinyue Jiang, Mohammed Bennamoun, Farid Boussaid, Xinyuan Qian, Qiuhong Ke·· 3 小时前
SAVU-Bench:面向空间音视频理解的真实世界基准
SAVU-BENCH: A Real-World Benchmark for Spatial Audio-Visual Understanding
AI 导读
研究者推出 SAVU-Bench,一个覆盖三个能力层级、七项评测任务的空间音视频理解真实世界基准,并配套 SAVU-Diag 诊断集,将推理问题拆解为前提性的定位与对齐子任务。对 12 个代表性模型的评测显示,视觉空间定位已相对成熟,涉及音频的空间感知仍是主要瓶颈。免训练的证据增强基线 SAVU-EA 显著提升空间定位与联合匹配,但高层空间推理依然困难。
正文
Abstract:Spatial audio-visual understanding requires models to recognize not only what is present, but also where events occur and how they relate across modalities. Existing benchmarks often rely on simulated scenes, evaluate isolated spatial skills, and provide limited diagnostic insight into failure modes. We introduce SAVU-Bench, a real-world benchmark that systematically evaluates spatial audio-visual understanding across three capability levels and seven evaluation tasks. We further introduce SAVU-Diag, a scene-linked diagnostic set that decomposes reasoning questions into their prerequisite grounding and alignment sub-tasks. Evaluation of 12 representative models on SAVU-Bench reveals that while visual spatial grounding is relatively mature, spatial perception involving audio remains a primary bottleneck. SAVU-Diag further demonstrates that most reasoning errors co-occur with failures on these prerequisite tasks, though reasoning gaps persist even when prerequisites are correctly resolved. Motivated by these findings, we introduce SAVU-EA, a training-free evidence-augmented baseline that makes spatial cues more explicit. While SAVU-EA substantially improves spatial grounding and joint matching, high-level spatial reasoning remains challenging. Our findings highlight the urgent need for both robust spatial audio perception and deeper integration of cross-modal spatial relations.
| Comments: | Under Review |
| Subjects: | Sound (cs.SD); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10624 [cs.SD] |
| (or arXiv:2610.10624v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10624 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yu Chen [view email]
[v1]
Wed, 7 Oct 2026 08:55:27 UTC (16,232 KB)
来源:arXiv:cs.AI · arxiv.org