TY - GEN
T1 - From Skeleton to Flesh
T2 - 2025 International Conference on Multimedia Retrieval, ICMR 2025
AU - Cao, Qianwen
AU - Huang, Heyan
AU - Wang, Boran
N1 - Publisher Copyright:
© 2025 ACM.
PY - 2025/6/30
Y1 - 2025/6/30
N2 - Video captioning is the task of automatically producing natural language descriptions for a given video. The great success of the Transformer architecture in NLP and CV has inspired recent attempts at constructing an end-to-end Transformer-based model for video captioning. Although achieving impressive performance, the one-step framework relies on brute force to learn highly compressed semantics from a large number of video patches, lacking the review process, which is a common behavior of human beings for understanding videos/papers/books. Taking the rough but significant impression captured at first sight as the skeleton, additional review can verify and investigate specific information as flesh to deepen precise recognition. In this work, we introduce the review process in the Transformer-based framework and propose a novel network, Aggregated Relational Transformer (ART) to conduct two-step decoding for video captioning. Since the relation triplet concisely summarizes the main structure, it is assigned as the objective of the first-pass decoding. Then the relations are utilized as a skeleton and guide the review process to potentially obtain better captions by looking into semantic components with a global view, where relations can also play the role of the prompt signal for the controllable caption generation at the second decoding pass. Extensive experiments show that our method achieves the SOTA performance on MSVD, MSRVTT, and VATEX datasets for video captioning, and is capable of controlling captions to respond to different semantic contexts.
AB - Video captioning is the task of automatically producing natural language descriptions for a given video. The great success of the Transformer architecture in NLP and CV has inspired recent attempts at constructing an end-to-end Transformer-based model for video captioning. Although achieving impressive performance, the one-step framework relies on brute force to learn highly compressed semantics from a large number of video patches, lacking the review process, which is a common behavior of human beings for understanding videos/papers/books. Taking the rough but significant impression captured at first sight as the skeleton, additional review can verify and investigate specific information as flesh to deepen precise recognition. In this work, we introduce the review process in the Transformer-based framework and propose a novel network, Aggregated Relational Transformer (ART) to conduct two-step decoding for video captioning. Since the relation triplet concisely summarizes the main structure, it is assigned as the objective of the first-pass decoding. Then the relations are utilized as a skeleton and guide the review process to potentially obtain better captions by looking into semantic components with a global view, where relations can also play the role of the prompt signal for the controllable caption generation at the second decoding pass. Extensive experiments show that our method achieves the SOTA performance on MSVD, MSRVTT, and VATEX datasets for video captioning, and is capable of controlling captions to respond to different semantic contexts.
KW - controllable text generation
KW - semantic bridging
KW - video captioning
UR - https://www.scopus.com/pages/publications/105011673070
U2 - 10.1145/3731715.3733347
DO - 10.1145/3731715.3733347
M3 - Conference contribution
AN - SCOPUS:105011673070
T3 - ICMR 2025 - Proceedings of the 2025 International Conference on Multimedia Retrieval
SP - 61
EP - 70
BT - ICMR 2025 - Proceedings of the 2025 International Conference on Multimedia Retrieval
PB - Association for Computing Machinery, Inc
Y2 - 30 June 2025 through 3 July 2025
ER -