arXiv:cs.AI· Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen·· 6 小时前AI 评分45
WorldSonus:为世界模型带来声音
WorldSonus: Bringing Sound to Worlds
AI 导读
WorldSonus 是一个面向世界模型的交互式视频转音频框架,可实时合成空间音频,实时因子低至 0.41。它采用流式因果自回归扩散架构,并通过以音频为中心的描述流水线与分块索引提示词调度,支持在生成过程中动态操控声音事件。基于多源立体声与 ambisonic 数据的高质量立体声监督,WorldSonus 在声学质量与空间对齐上匹配或超越当前双向模型。
正文
Authors:Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen
Abstract:Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: this https URL
| Comments: | 25 pages, 4 figures, 16 tables. Project page: this https URL |
| Subjects: | Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.08760 [cs.SD] |
| (or arXiv:2610.08760v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08760 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pengjun Fang [view email]
[v1]
Tue, 6 Oct 2026 17:50:18 UTC (10,428 KB)
来源:arXiv:cs.AI · arxiv.org