Skip to main navigation Skip to search Skip to main content

Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning

  • Rongjiang Zhu
  • , Xinxiao Wu
  • , Shuo Yang*
  • , Yuheng Shi
  • , Ziyi Wang
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University

Research output: Contribution to journalArticlepeer-review

Abstract

Open-vocabulary multi-label action recognition is extremely challenging since it requires identifying a broader spectrum of multiple unseen actions during inference. Recent progresses have been made by designing or learning textual prompts based on action class labels to adapt pre-trained Vision-Language Models (VLMs), but neglecting co-occurrence relationships among actions in multi-label scenarios, limiting the performance. We propose a Large Language Model (LLM) enhanced prompt tuning method, which uses commonsense knowledge of action dependencies extracted from LLMs. By posing carefully designed queries, we obtain rich descriptions of co-occurrence relationships between actions, which are disentangled into temporal and spatial dependencies to capture distinct interaction types. These insights are then integrated through learnable correlation prompts, enabling better adaptation to complex multi-action scenarios. Additionally, we manually annotate action videos from the MovieNet dataset using multiple class labels to evaluate our method. Extensive experiments demonstrate that our method achieves better results than state-of-the-art methods. Source codes and the newly annotated dataset will be available.

Original languageEnglish
Article number104876
JournalComputer Vision and Image Understanding
Volume271
DOIs
Publication statusPublished - Sept 2026
Externally publishedYes

Keywords

  • Large language model (LLM)
  • Open-vocabulary action recognition
  • Prompt learning

Fingerprint

Dive into the research topics of 'Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning'. Together they form a unique fingerprint.

Cite this