跳到主要导航 跳到搜索 跳到主要内容

VidEvo: Evolving Video Editing through Exhaustive Temporal Modeling

  • Sizhe Dang
  • , Huan Liu
  • , Mengmeng Wang
  • , Xin Lai
  • , Guang Dai
  • , Jingdong Wang
  • Xi'an Jiaotong University
  • Zhejiang University of Technology
  • State Grid Corporation of China
  • Baidu Inc

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Text-guided video editing (TGVE) has become a recent hotspot due to its entertainment value and practical applications. To reduce overhead, existing methods primarily extend from text-to-image diffusion models and typically involve reconstruction and editing phases. However, challenges persist, particularly in enhancing temporal consistency of a video while adhering to textual alignment requirements. A crucial factor leading to the aforementioned issue is the inadequate and implicit tuning of the attention module within existing methods, which is specifically designed to capture temporal information. In light of this, we introduce VidEvo, a novel one-shot video editing method that leverages explicit cues derived from the original video to enhance temporal modeling. By integrating null-video embedding (NVE) and window-frame attention (WFA) components, VidEvo facilitates the smooth and coherent generation of videos from global and local perspectives simultaneously. To be specific, NVE learns a set of multi-scale temporal embeddings within the visual space during the reconstruction phase. These embeddings are subsequently directly injected into the attention module of the editing phase, explicitly augmenting the temporal consistency of the entire video. On the other hand, WFA enhances local temporal modeling by dynamically optimizing attention mechanisms between adjacent frames, which improves temporal coherence with reduced computational costs. Experimental evaluations show that VidEvo enhances frame-to-frame temporal consistency. Ablation studies confirm NVE and WFA's effectiveness and their plug-and-play capability with other methods.

源语言英语
主期刊名Proceedings of the 34th International Joint Conference on Artificial Intelligence, IJCAI 2025
编辑James Kwok
出版商International Joint Conferences on Artificial Intelligence
882-890
页数9
ISBN(电子版)9781956792065
DOI
出版状态已出版 - 2025
活动34th Internationa Joint Conference on Artificial Intelligence, IJCAI 2025 - Montreal, 加拿大
期限: 16 8月 202522 8月 2025

丛书

姓名IJCAI International Joint Conference on Artificial Intelligence
ISSN(印刷版)1045-0823

会议

会议34th Internationa Joint Conference on Artificial Intelligence, IJCAI 2025
国家/地区加拿大
Montreal
时期16/08/2522/08/25

学术指纹

探究 'VidEvo: Evolving Video Editing through Exhaustive Temporal Modeling' 的科研主题。它们共同构成独一无二的学术指纹。

引用此