arXiv:cs.LG· Wenhao Sun, Rong-Cheng Tu, Jingyi Liao, Dacheng Tao·· 6 小时前AI 评分30
基于扩散模型的视频编辑综述:方法、应用与 V2VBench 基准
Diffusion Model-Based Video Editing: A Survey
AI 导读
这篇综述系统梳理了基于扩散模型的视频编辑技术,涵盖其数学基础与图像域关键方法,并按核心技术的内在联系对视频编辑方法进行分类,勾勒出技术演进脉络。文章还探讨了点编辑、姿态引导的人体视频编辑等新应用,并基于新提出的 V2VBench 进行了全面对比。全文 24 页、16 幅图,最后总结了当前挑战与未来研究方向。
正文
Abstract:The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality. Among these, video editing has gained substantial attention and seen a swift rise in research activity, necessitating a comprehensive and systematic review of the existing literature. This paper reviews diffusion model-based video editing techniques, including theoretical foundations and practical applications. We begin by overviewing the mathematical formulation and image domain's key methods. Subsequently, we categorize video editing approaches by the inherent connections of their core technologies, depicting evolutionary trajectory. This paper also dives into novel applications, including point-based editing and pose-guided human video editing. Additionally, we present a comprehensive comparison using our newly introduced V2VBench. Building on the progress achieved to date, the paper concludes with ongoing challenges and potential directions for future research.
| Comments: | 24 pages, 16 figures, a project related to this paper can be found at this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM) |
| Cite as: | arXiv:2407.07111 [cs.CV] |
| (or arXiv:2407.07111v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2407.07111 arXiv-issued DOI via DataCite |
|
| Journal reference: | International Journal of Computer Vision 134(10), 440 (2026) |
| Related DOI: | https://doi.org/10.1007/s11263-026-03040-6
DOI(s) linking to related resources |
Submission history
From: Wenhao Sun [view email]
[v1]
Wed, 26 Jun 2024 04:58:39 UTC (3,577 KB)
[v2]
Tue, 6 Oct 2026 08:18:15 UTC (14,314 KB)
来源:arXiv:cs.LG · arxiv.org