跳到主要导航 跳到搜索 跳到主要内容

Jigsaw-ViT: Learning jigsaw puzzles in vision transformer

  • Yingyi Chen*
  • , Xi Shen
  • , Yahui Liu
  • , Qinghua Tao
  • , Johan A.K. Suykens
  • *此作品的通讯作者
  • KU Leuven
  • Tencent
  • University of Trento

科研成果: 期刊稿件文章同行评审

摘要

The success of Vision Transformer (ViT) in various computer vision tasks has promoted the ever-increasing prevalence of this convolution-free network. The fact that ViT works on image patches makes it potentially relevant to the problem of jigsaw puzzle solving, which is a classical self-supervised task aiming at reordering shuffled sequential image patches back to their original form. Solving jigsaw puzzle has been demonstrated to be helpful for diverse tasks using Convolutional Neural Networks (CNNs), such as feature representation learning, domain generalization and fine-grained classification. In this paper, we explore solving jigsaw puzzle as a self-supervised auxiliary loss in ViT for image classification, named Jigsaw-ViT. We show two modifications that can make Jigsaw-ViT superior to standard ViT: discarding positional embeddings and masking patches randomly. Yet simple, we find that the proposed Jigsaw-ViT is able to improve on both generalization and robustness over the standard ViT, which is usually rather a trade-off. Numerical experiments verify that adding the jigsaw puzzle branch provides better generalization to ViT on large-scale image classification on ImageNet. Moreover, such auxiliary loss also improves robustness against noisy labels on Animal-10N, Food-101N, and Clothing1M, as well as adversarial examples. Our implementation is available at https://yingyichen-cyy.github.io/Jigsaw-ViT.

源语言英语
页(从-至)53-60
页数8
期刊Pattern Recognition Letters
166
DOI
出版状态已出版 - 2月 2023
已对外发布

学术指纹

探究 'Jigsaw-ViT: Learning jigsaw puzzles in vision transformer' 的科研主题。它们共同构成独一无二的学术指纹。

引用此