arXiv:cs.LG· Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang·· 5 小时前AI 评分34
VIDiff:用扩散模型和多模态指令实现视频翻译
VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models
AI 导读
研究者提出统一基础模型 Video Instruction Diffusion(VIDiff),可基于用户指令在数秒内完成视频编辑与翻译,并覆盖语言引导的视频目标分割等理解任务。该模型采用迭代自回归方法保证长视频编辑与增强的一致性,在多种输入视频和文本指令上均取得有说服力的生成结果。
正文
Abstract:Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches only focus on video editing for short clips and rely on time-consuming tuning or inference. We are the first to propose Video Instruction Diffusion (VIDiff), a unified foundation model designed for a wide range of video tasks. These tasks encompass both understanding tasks (such as language-guided video object segmentation) and generative tasks (video editing and enhancement). Our model can edit and translate the desired results within seconds based on user instructions. Moreover, we design an iterative auto-regressive method to ensure consistency in editing and enhancing long videos. We provide convincing generative results for diverse input videos and written instructions, both qualitatively and quantitatively. More examples can be found at our website this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM) |
| Cite as: | arXiv:2311.18837 [cs.CV] |
| (or arXiv:2311.18837v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2311.18837 arXiv-issued DOI via DataCite |
Submission history
From: Zhen Xing [view email]
[v1]
Thu, 30 Nov 2023 18:59:52 UTC (21,056 KB)
[v2]
Fri, 2 Oct 2026 08:56:54 UTC (21,042 KB)
来源:arXiv:cs.LG · arxiv.org