跳到主要导航 跳到搜索 跳到主要内容

Video semantic segmentation via feature propagation with holistic attention

  • Junrong Wu
  • , Zongzheng Wen
  • , Sanyuan Zhao*
  • , Kele Huang
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • University of Chinese Academy of Sciences

科研成果: 期刊稿件文章同行评审

摘要

Since the frames of a video are inherently contiguous, information redundancy is ubiquitous. Unlike previous works densely process each frame of a video, in this paper we present a novel method to focus on efficient feature propagation across frames to tackle the challenging video semantic segmentation task. Firstly, we propose a Light, Efficient and Real-time network (denoted as LERNet) as a strong backbone network for per-frame processing. Then we mine rich features within a key frame and propagate the across-frame consistency information by calculating a temporal holistic attention with the following non-key frame. Each element of the attention matrix represents the global correlation between pixels of a non-key frame and the previous key frame. Concretely, we propose a brand-new attention module to capture the spatial consistency on low-level features along temporal dimension. Then we employ the attention weights as a spatial transition guidance for directly generating high-level features of the current non-key frame from the weighted corresponding key frame. Finally, we efficiently fuse the hierarchical features of the non-key frame and obtain the final segmentation result. Extensive experiments on two popular datasets, i.e. the CityScapes and the CamVid, demonstrate that the proposed approach achieves a remarkable balance between inference speed and accuracy.

源语言英语
文章编号107268
期刊Pattern Recognition
104
DOI
出版状态已出版 - 8月 2020

学术指纹

探究 'Video semantic segmentation via feature propagation with holistic attention' 的科研主题。它们共同构成独一无二的学术指纹。

引用此