跳到主要导航 跳到搜索 跳到主要内容

HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation

  • Boyuan Wang
  • , Xiaofeng Wang
  • , Chaojun Ni
  • , Guosheng Zhao
  • , Zhiqin Yang
  • , Zheng Zhu*
  • , Muyang Zhang
  • , Yukun Zhou
  • , Xinze Chen
  • , Guan Huang
  • , Lihong Liu
  • , Xingang Wang*
  • *此作品的通讯作者
  • CAS - Institute of Automation
  • University of Chinese Academy of Sciences
  • Luoyang Institute for Robot and Intelligent Equipment
  • GigaAI
  • Peking University
  • Chinese University of Hong Kong

科研成果: 期刊稿件会议文章同行评审

摘要

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from existing videos, thereby lacking flexibility. To address this, we propose HumanDreamer, a decoupled human video generation framework that first generates diverse poses from text prompts and then leverages these poses to generate human-motion videos. Specifically, we propose MotionVid, the largest dataset for human-motion pose generation. Based on the dataset, we present MotionDiT, which is trained to generate structured human-motion poses from text prompts. Besides, a novel LAMA loss is introduced, which together contribute to a significant improvement in FID by 62.4%, along with respective enhancements in R-precision for top1, top2, and top3 by 41.8%, 26.3%, and 18.3%, thereby advancing both the Text-to-Pose control accuracy and FID metrics. Our experiments across various Pose-to-Video baselines demonstrate that the poses generated by our method can produce diverse and high-quality human-motion videos. Furthermore, our model can facilitate other downstream tasks, such as pose sequence prediction and 2D-3D motion lifting.

源语言英语
页(从-至)12391-12401
页数11
期刊Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOI
出版状态已出版 - 2025
已对外发布
活动2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, 美国
期限: 11 6月 202515 6月 2025

学术指纹

探究 'HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此