跳到主要导航 跳到搜索 跳到主要内容

Commonsense Knowledge Prompting for Few-Shot Action Recognition in Videos

  • Yuheng Shi
  • , Xinxiao Wu*
  • , Hanxi Lin
  • , Jiebo Luo
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University
  • University of Rochester

科研成果: 期刊稿件文章同行评审

摘要

—Few-shot action recognition in videos is challenging as the lack of supervision makes it extremely difficult to generalize well to unseen actions. To address this challenge, we propose a simple yet effective method, called knowledge prompting, which leverages commonsense knowledge of actions from external resources to prompt-tune a powerful pre-trained vision-language model for few-shot classification. To that end, we first collect a large-scale corpus of language descriptions of actions, defined as text proposals, to build an action knowledge base. The collection of text proposals is done by filling in a handcraft sentence template with an external action-related corpus or by extracting action-related phrases from captions of Web instruction videos. Next, we feed these text proposals to a pre-trained vision-language model along with video frames to generate matching scores of the proposals for each frame, and the scores can be treated as action semantics with strong generalization. Finally, we design a lightweight temporal modeling network to capture the temporal evolution of action semantics for classification. Extensive experiments on six benchmark datasets demonstrate that our method generally achieves state-of-the-art performance while reducing the training computational cost to 0.1% of the existing methods.

源语言英语
页(从-至)8395-8405
页数11
期刊IEEE Transactions on Multimedia
26
DOI
出版状态已出版 - 2024
已对外发布

学术指纹

探究 'Commonsense Knowledge Prompting for Few-Shot Action Recognition in Videos' 的科研主题。它们共同构成独一无二的学术指纹。

引用此