视频-音频联合与跨模态生成及编辑:统一形式化与设计分类体系
Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
一篇综述提出将联合生成、跨模态生成与联合编辑统一建模为音视频对上的同一分布问题,并按五个设计轴对方法进行分类。作者称这是首个系统梳理联合音视频编辑的综述,将其划分为九类编辑、涵盖 28 种编辑类型,并整理了各场景的方法、数据集与评测指标。
Published on Sep 28
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
View arXiv page View PDF Add to collection
Get this paper in your agent:
hf papers read 2609.34381
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.34381 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.34381 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.34381 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co