跳到主要导航 跳到搜索 跳到主要内容

DriveGen: Shared Video-Condition Encoding for Autonomous Multi-View Video Generation

  • Yuhao Kang
  • , Haolin Li
  • , Sanyuan Zhao*
  • , Shudong Wang
  • , Xiameng Qin
  • , Junyu Han
  • , Ji Tao
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Changan Automobile

科研成果: 期刊稿件文章同行评审

摘要

Corner cases, such as severe weather and abnormal lighting, present significant challenges in autonomous driving. The main obstacles involve large-scale data collection and costly annotations. Leveraging generative models to expand corner-case data based on existing annotations offers a promising solution. Unlike monocular videos, multi-view videos introduce an additional “view” dimension, increasing the consistency requirements and making precise control of annotations more challenging. Existing methods decouple multi-view videos along the temporal and view-spatial axes, using separate attention mechanisms, which causes motion discrepancies and limits consistency. Additionally, current approaches employ an independent adapter or ControlNet to encode different 3D annotations, leading to high computational costs and suboptimal alignment between annotations and video latents. These issues arise from neglecting the temporal-spatial relationship and insufficient alignment between 3D annotations and video latents. To address these challenges, we propose DriveGen, which uses 4D position embeddings to encode the positional information of multi-view videos. DriveGen also designs Dual-Scale Full Attention to ensure both global and local spatiotemporal consistency. Furthermore, our Shared Video-Condition Encoding (SVCE) Mechanism converts 3D annotations into 2D masks and encodes both video and annotation sequences using a 3D VAE, requiring only 0.37 M learnable parameters to achieve pixel-level alignment and improving generation quality. Numerous experiments have proven that DriveGen has reached the state-of-the-art, capable of generating high-quality controlled autonomous driving videos.

源语言英语
期刊IEEE Transactions on Visualization and Computer Graphics
DOI
出版状态已接受/待刊 - 2026
已对外发布

指纹

探究 'DriveGen: Shared Video-Condition Encoding for Autonomous Multi-View Video Generation' 的科研主题。它们共同构成独一无二的指纹。

引用此