Skip to main navigation Skip to search Skip to main content

From Skeleton to Flesh: Aggregated Relational Transformer Towards Controllable Video Captioning with Two-Step Decoding

  • China University of Petroleum - Beijing
  • ByteDance Ltd.

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Video captioning is the task of automatically producing natural language descriptions for a given video. The great success of the Transformer architecture in NLP and CV has inspired recent attempts at constructing an end-to-end Transformer-based model for video captioning. Although achieving impressive performance, the one-step framework relies on brute force to learn highly compressed semantics from a large number of video patches, lacking the review process, which is a common behavior of human beings for understanding videos/papers/books. Taking the rough but significant impression captured at first sight as the skeleton, additional review can verify and investigate specific information as flesh to deepen precise recognition. In this work, we introduce the review process in the Transformer-based framework and propose a novel network, Aggregated Relational Transformer (ART) to conduct two-step decoding for video captioning. Since the relation triplet concisely summarizes the main structure, it is assigned as the objective of the first-pass decoding. Then the relations are utilized as a skeleton and guide the review process to potentially obtain better captions by looking into semantic components with a global view, where relations can also play the role of the prompt signal for the controllable caption generation at the second decoding pass. Extensive experiments show that our method achieves the SOTA performance on MSVD, MSRVTT, and VATEX datasets for video captioning, and is capable of controlling captions to respond to different semantic contexts.

Original languageEnglish
Title of host publicationICMR 2025 - Proceedings of the 2025 International Conference on Multimedia Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages61-70
Number of pages10
ISBN (Electronic)9798400718779
DOIs
Publication statusPublished - 30 Jun 2025
Event2025 International Conference on Multimedia Retrieval, ICMR 2025 - Chicago, United States
Duration: 30 Jun 20253 Jul 2025

Publication series

NameICMR 2025 - Proceedings of the 2025 International Conference on Multimedia Retrieval

Conference

Conference2025 International Conference on Multimedia Retrieval, ICMR 2025
Country/TerritoryUnited States
CityChicago
Period30/06/253/07/25

Keywords

  • controllable text generation
  • semantic bridging
  • video captioning

Fingerprint

Dive into the research topics of 'From Skeleton to Flesh: Aggregated Relational Transformer Towards Controllable Video Captioning with Two-Step Decoding'. Together they form a unique fingerprint.

Cite this