arXiv:cs.LG· Kyeong-Rae Kim, Sungnyun Kim, Tae-Hyun Oh·· 3 小时前
FloorSAV:用 2D 平面图为 AV-LLM 注入空间音视频上下文
FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
AI 导读
FloorSAV 通过渲染动态 2D floormap,将 3D 点云、相机轨迹、空间音频线索与语义物体地标融合为与第一人称视频同步的流输入 AV-LLM,让模型在单次推理中联合理解视觉、听觉与几何线索。该框架提升了 AV-LLM 在 SAVED-Bench 与 SAVVY-Bench 多项空间推理任务上的表现,并引入包含动态相对位置、区域和路径推理 QA 的新基准 SAVED-Bench。
正文
Abstract:While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
| Comments: | Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11310 [cs.CV] |
| (or arXiv:2610.11310v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11310 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gyeongrae Kim [view email]
[v1]
Thu, 8 Oct 2026 06:15:36 UTC (12,737 KB)
来源:arXiv:cs.LG · arxiv.org