跳到正文
HuggingFace Daily Papers·· 1 天前AI 评分36

ALIVE:面向首帧引导视频编辑的交互对齐物体插入框架

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

AI 导读

ALIVE 是一个让插入物体与源视频内容产生连贯交互的视频编辑框架,仅需一张编辑后的首帧和一条只命名新增物体的指令。团队构建了 35,800 对编辑数据,并训练 VLM 预测交互引导;在无 VLM 引导下,ALIVE 在两个基准上的 Overall 指标分别比最强基线提升 43.9% 和 4.4%,加入 VLM 预测引导后 ALIVE-interaction 分数再提升 0.95 分。

正文

View PDF HTML (experimental)

Abstract:Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects "alive" through coherent interactions with the source video's contents, using an edited first frame and an instruction naming only the added object. We curate 35,800 editing pairs combining 3D-rendered, model-generated, and real-world videos with general editing pairs from ROSE. Each pair differs in the target object's presence while preserving the surrounding action, teaching editors coordinated object behavior and source preservation. We further train a vision-language model (VLM) to predict interaction guidance from the same inputs. We introduce the ALIVE-interaction benchmark to assess interaction fidelity, source preservation, and visual coherence using a unified VLM-based protocol, and evaluate on the general video object insertion benchmark. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% and 4.4% on the two benchmarks, respectively. VLM-predicted guidance further improves the ALIVE-interaction score by 0.95 points without additional user inputs.
Comments: Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.08779 [cs.CV]
  (or arXiv:2610.08779v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.08779

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhenghong Zhou [view email]
[v1] Tue, 6 Oct 2026 17:58:51 UTC (43,045 KB)

来源:HuggingFace Daily Papers · arxiv.org