Topic-aware video summarization using multimodal transformer

Yubo Zhu, Wentian Zhao, Rui Hua, Xinxiao Wu*

*此作品的通讯作者

科研成果: 期刊稿件文章同行评审

10 引用 (Scopus)

摘要

Video summarization aims to generate a short and compact summary to represent the original video. Existing methods mainly focus on how to extract a general objective synopsis that precisely summaries the video content. However, in real scenarios, a video usually contains rich content with multiple topics and people may cast diverse interests on the visual contents even for the same video. In this paper, we propose a novel topic-aware video summarization task that generates multiple video summaries with different topics. To support the study of this new task, we first build a video benchmark dataset by collecting videos from various types of movies and annotate them with topic labels and frame-level importance scores. Then we propose a multimodal Transformer model for the topic-aware video summarization, which simultaneously predicts topic labels and generates topic-related summaries by adaptively fusing multimodal features extracted from the video. Experimental results show the effectiveness of our method.

源语言英语
文章编号109578
期刊Pattern Recognition
140
DOI
出版状态已出版 - 8月 2023
已对外发布

指纹

探究 'Topic-aware video summarization using multimodal transformer' 的科研主题。它们共同构成独一无二的指纹。

引用此