Abstract
Open-vocabulary multi-label action recognition is extremely challenging since it requires identifying a broader spectrum of multiple unseen actions during inference. Recent progresses have been made by designing or learning textual prompts based on action class labels to adapt pre-trained Vision-Language Models (VLMs), but neglecting co-occurrence relationships among actions in multi-label scenarios, limiting the performance. We propose a Large Language Model (LLM) enhanced prompt tuning method, which uses commonsense knowledge of action dependencies extracted from LLMs. By posing carefully designed queries, we obtain rich descriptions of co-occurrence relationships between actions, which are disentangled into temporal and spatial dependencies to capture distinct interaction types. These insights are then integrated through learnable correlation prompts, enabling better adaptation to complex multi-action scenarios. Additionally, we manually annotate action videos from the MovieNet dataset using multiple class labels to evaluate our method. Extensive experiments demonstrate that our method achieves better results than state-of-the-art methods. Source codes and the newly annotated dataset will be available.
| Original language | English |
|---|---|
| Article number | 104876 |
| Journal | Computer Vision and Image Understanding |
| Volume | 271 |
| DOIs | |
| Publication status | Published - Sept 2026 |
| Externally published | Yes |
Keywords
- Large language model (LLM)
- Open-vocabulary action recognition
- Prompt learning
Fingerprint
Dive into the research topics of 'Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver