跳到主要导航 跳到搜索 跳到主要内容

Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning

  • Xiangyan Qu
  • , Jing Yu*
  • , Keke Gai
  • , Jiamin Zhuang
  • , Yuanmin Tang
  • , Gang Xiong
  • , Gaopeng Gou
  • , Qi Wu
  • *此作品的通讯作者
  • CAS - Institute of Information Engineering
  • Beijing Institute of Technology
  • Adelaide University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with corresponding images to transfer knowledge. However, they disregard that semantic information is not equivalent between them, resulting in a suboptimal alignment. In this work, we propose a novel network to extract multi-view semantic concepts from documents and images and align the matching rather than entire concepts. Specifically, we propose a semantic decomposition module to generate multi-view semantic embeddings from visual and textual sides, providing the basic concepts for partial alignment. To alleviate the issue of information redundancy among embeddings, we propose the local-to-semantic variance loss to capture distinct local details and multiple semantic diversity loss to enforce orthogonality among embeddings. Subsequently, two losses are introduced to partially align visual-semantic embedding pairs according to their semantic relevance at the view and word-to-patch levels. Consequently, we consistently outperform state-of-the-art methods under two document sources in three standard benchmarks for document-based zero-shot learning. Qualitatively, we show that our model learns the interpretable partial association. Code is available at https://github.com/MorningStarOvO/EmDepart.

源语言英语
主期刊名MM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia
出版商Association for Computing Machinery, Inc
4581-4590
页数10
ISBN(电子版)9798400706868
DOI
出版状态已出版 - 28 10月 2024
已对外发布
活动32nd ACM International Conference on Multimedia, MM 2024 - Melbourne, 澳大利亚
期限: 28 10月 20241 11月 2024

出版系列

姓名MM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia

会议

会议32nd ACM International Conference on Multimedia, MM 2024
国家/地区澳大利亚
Melbourne
时期28/10/241/11/24

指纹

探究 'Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning' 的科研主题。它们共同构成独一无二的指纹。

引用此