跳到主要导航 跳到搜索 跳到主要内容

Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning

  • Rongjiang Zhu
  • , Xinxiao Wu
  • , Shuo Yang*
  • , Yuheng Shi
  • , Ziyi Wang
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University

科研成果: 期刊稿件文章同行评审

摘要

Open-vocabulary multi-label action recognition is extremely challenging since it requires identifying a broader spectrum of multiple unseen actions during inference. Recent progresses have been made by designing or learning textual prompts based on action class labels to adapt pre-trained Vision-Language Models (VLMs), but neglecting co-occurrence relationships among actions in multi-label scenarios, limiting the performance. We propose a Large Language Model (LLM) enhanced prompt tuning method, which uses commonsense knowledge of action dependencies extracted from LLMs. By posing carefully designed queries, we obtain rich descriptions of co-occurrence relationships between actions, which are disentangled into temporal and spatial dependencies to capture distinct interaction types. These insights are then integrated through learnable correlation prompts, enabling better adaptation to complex multi-action scenarios. Additionally, we manually annotate action videos from the MovieNet dataset using multiple class labels to evaluate our method. Extensive experiments demonstrate that our method achieves better results than state-of-the-art methods. Source codes and the newly annotated dataset will be available.

源语言英语
期刊论文编号104876
期刊Computer Vision and Image Understanding
271
DOI
出版状态已出版 - 9月 2026
已对外发布

学术指纹

探究 'Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此