跳到主要导航 跳到搜索 跳到主要内容

Adversarial task-specific learning

  • Xin Fu
  • , Yao Zhao*
  • , Ting Liu
  • , Yunchao Wei
  • , Jianan Li
  • , Shikui Wei
  • *此作品的通讯作者
  • Beijing Jiaotong University
  • University of Illinois at Urbana-Champaign

科研成果: 期刊稿件文章同行评审

摘要

In this paper, we investigate a principle way to learn a common feature space for data of different modalities (e.g. image and text), so that the similarity between different modal items can be directly measured for benefiting cross-modal retrieval task. To effectively keep semantic/distribution consistent for common feature embeddings, we propose a new Adversarial Task-Specific Learning (ATSL) approach to learn distinct embeddings for different retrieval tasks, i.e. images retrieve texts (I2T) or texts retrieve images (T2I). In particular, the proposed ATSL is with the following advantages: (a) semantic attributes are leveraged to encourage the learned common feature embeddings of couples to be semantic consistent; (b) adversarial learning is applied to relieve the inconsistent distribution of common feature embeddings for different modalities; (c) triplet optimization is employed to guarantee that similar items from different modalities are with smaller distances in the learned common space compared with the dissimilar ones; (d) task-specific learning produces better optimized common feature embeddings for different retrieval tasks. Our ATSL is embedded in a deep neural network, which can be learned in an end-to-end manner. We conduct extensive experiments on two popular benchmark datasets, e.g. Flickr30K and MS COCO. We achieve R@1 accuracy of 57.1% and 38.4% for I2T and 56.5% and 38.6% T2I on MS COCO and Flickr30K respectively, which are the new state-of-the-arts.

源语言英语
页(从-至)118-128
页数11
期刊Neurocomputing
362
DOI
出版状态已出版 - 14 10月 2019

学术指纹

探究 'Adversarial task-specific learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此