TY - JOUR
T1 - Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning
AU - Zhu, Rongjiang
AU - Wu, Xinxiao
AU - Yang, Shuo
AU - Shi, Yuheng
AU - Wang, Ziyi
N1 - Publisher Copyright:
© 2026 Elsevier Inc.
PY - 2026/9
Y1 - 2026/9
N2 - Open-vocabulary multi-label action recognition is extremely challenging since it requires identifying a broader spectrum of multiple unseen actions during inference. Recent progresses have been made by designing or learning textual prompts based on action class labels to adapt pre-trained Vision-Language Models (VLMs), but neglecting co-occurrence relationships among actions in multi-label scenarios, limiting the performance. We propose a Large Language Model (LLM) enhanced prompt tuning method, which uses commonsense knowledge of action dependencies extracted from LLMs. By posing carefully designed queries, we obtain rich descriptions of co-occurrence relationships between actions, which are disentangled into temporal and spatial dependencies to capture distinct interaction types. These insights are then integrated through learnable correlation prompts, enabling better adaptation to complex multi-action scenarios. Additionally, we manually annotate action videos from the MovieNet dataset using multiple class labels to evaluate our method. Extensive experiments demonstrate that our method achieves better results than state-of-the-art methods. Source codes and the newly annotated dataset will be available.
AB - Open-vocabulary multi-label action recognition is extremely challenging since it requires identifying a broader spectrum of multiple unseen actions during inference. Recent progresses have been made by designing or learning textual prompts based on action class labels to adapt pre-trained Vision-Language Models (VLMs), but neglecting co-occurrence relationships among actions in multi-label scenarios, limiting the performance. We propose a Large Language Model (LLM) enhanced prompt tuning method, which uses commonsense knowledge of action dependencies extracted from LLMs. By posing carefully designed queries, we obtain rich descriptions of co-occurrence relationships between actions, which are disentangled into temporal and spatial dependencies to capture distinct interaction types. These insights are then integrated through learnable correlation prompts, enabling better adaptation to complex multi-action scenarios. Additionally, we manually annotate action videos from the MovieNet dataset using multiple class labels to evaluate our method. Extensive experiments demonstrate that our method achieves better results than state-of-the-art methods. Source codes and the newly annotated dataset will be available.
KW - Large language model (LLM)
KW - Open-vocabulary action recognition
KW - Prompt learning
UR - https://www.scopus.com/pages/publications/105044996363
U2 - 10.1016/j.cviu.2026.104876
DO - 10.1016/j.cviu.2026.104876
M3 - Article
AN - SCOPUS:105044996363
SN - 1077-3142
VL - 271
JO - Computer Vision and Image Understanding
JF - Computer Vision and Image Understanding
M1 - 104876
ER -