跳到正文
arXiv:cs.AI· Leigang Qu, Feng Cheng, Ziyan Yang, Bangbang Yang, Zhaoyang Huang, Wei Chow, Yicong Li, Wenjie Wang, Tat-Seng Chua, Yan Zeng·· 4 小时前

VINCIE-NExT:通过上下文建模从图像解锁视频编辑

VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling

AI 导读

VINCIE-NExT 通过上下文视觉演示将图像编辑能力迁移到视频,减少对大规模配对视频编辑数据的依赖。该框架将视频编辑分解为 Video → Image → Image → Video 子任务链,并引入共享空间坐标系的位置编码,将图像演示与视频帧关联,实现外观改动的逐像素传播。在 OpenVE-Bench 上取得 SOTA 表现,Chain-of-Editing 支持无需重训的测试时扩展。

正文

View PDF HTML (experimental)

Abstract:Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
Comments: Accepted to NeurIPS'26. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
Cite as: arXiv:2610.12104 [cs.CV]
  (or arXiv:2610.12104v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.12104

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Leigang Qu [view email]
[v1] Thu, 8 Oct 2026 15:03:48 UTC (20,286 KB)

来源:arXiv:cs.AI · arxiv.org