跳到主要导航 跳到搜索 跳到主要内容

TwinNet: Twin Structured Knowledge Transfer Network for Weakly Supervised Action Localization

  • Xiao Yu Zhang
  • , Hai Chao Shi*
  • , Chang Sheng Li
  • , Li Xin Duan
  • *此作品的通讯作者
  • CAS - Institute of Information Engineering
  • University of Electronic Science and Technology of China

科研成果: 期刊稿件文章同行评审

摘要

Action recognition and localization in untrimmed videos is important for many applications and have attracted a lot of attention. Since full supervision with frame-level annotation places an overwhelming burden on manual labeling effort, learning with weak video-level supervision becomes a potential solution. In this paper, we propose a novel weakly supervised framework to recognize actions and locate the corresponding frames in untrimmed videos simultaneously. Considering that there are abundant trimmed videos publicly available and well-segmented with semantic descriptions, the instructive knowledge learned on trimmed videos can be fully leveraged to analyze untrimmed videos. We present an effective knowledge transfer strategy based on inter-class semantic relevance. We also take advantage of the self-attention mechanism to obtain a compact video representation, such that the influence of background frames can be effectively eliminated. A learning architecture is designed with twin networks for trimmed and untrimmed videos, to facilitate transferable self-attentive representation learning. Extensive experiments are conducted on three untrimmed benchmark datasets (i.e., THUMOS14, ActivityNet1.3, and MEXaction2), and the experimental results clearly corroborate the efficacy of our method. It is especially encouraging to see that the proposed weakly supervised method even achieves comparable results to some fully supervised methods.

源语言英语
页(从-至)227-246
页数20
期刊Machine Intelligence Research
19
3
DOI
出版状态已出版 - 6月 2022

学术指纹

探究 'TwinNet: Twin Structured Knowledge Transfer Network for Weakly Supervised Action Localization' 的科研主题。它们共同构成独一无二的学术指纹。

引用此