Visual tracking using transformer with a combination of convolution and attention

Yuxuan Wang; Liping Yan; Zihang Feng; Yuanqing Xia; Bo Xiao

doi:10.1016/j.imavis.2023.104760

Visual tracking using transformer with a combination of convolution and attention

Yuxuan Wang, Liping Yan^*, Zihang Feng, Yuanqing Xia, Bo Xiao

^*Corresponding author for this work

School of Automation

Research output: Contribution to journal › Article › peer-review

Abstract

For Siamese-based trackers in the field of single object tracking, cross-correlation operation plays an important role. However, the cross-correlation essentially uses target feature to locally linearly match the search region, which leads to insufficient utilization or even loss of feature information. To effectively employ global context and sufficiently explore the relevance of template and search region, a novel matching operator is designed inspired by Transformer, which uses multi-head attention and embed a designed modulation module across the inputs of operator. Meanwhile, we equip our tracker with a multi-scale encoder/decoder strategy to gradually make more precise tracking. Finally, a complete tracking framework is presented named VTTR. The tracker consists of a feature extractor, a multi-scale encoder based on depth-wise convolution, a modified decoder as the matching operator and a prediction head. The proposed tracker is tested on many benchmarks and achieve excellent performance while running with fast speed.

Original language	English
Article number	104760
Journal	Image and Vision Computing
Volume	137
DOIs	https://doi.org/10.1016/j.imavis.2023.104760
Publication status	Published - Sept 2023

Keywords

Attention
Siamese networks
Transformer
Visual tracking

Access to Document

10.1016/j.imavis.2023.104760

Cite this

@article{bea3e4ee7db74b72956ea6e848e62b5c,

title = "Visual tracking using transformer with a combination of convolution and attention",

abstract = "For Siamese-based trackers in the field of single object tracking, cross-correlation operation plays an important role. However, the cross-correlation essentially uses target feature to locally linearly match the search region, which leads to insufficient utilization or even loss of feature information. To effectively employ global context and sufficiently explore the relevance of template and search region, a novel matching operator is designed inspired by Transformer, which uses multi-head attention and embed a designed modulation module across the inputs of operator. Meanwhile, we equip our tracker with a multi-scale encoder/decoder strategy to gradually make more precise tracking. Finally, a complete tracking framework is presented named VTTR. The tracker consists of a feature extractor, a multi-scale encoder based on depth-wise convolution, a modified decoder as the matching operator and a prediction head. The proposed tracker is tested on many benchmarks and achieve excellent performance while running with fast speed.",

keywords = "Attention, Siamese networks, Transformer, Visual tracking",

author = "Yuxuan Wang and Liping Yan and Zihang Feng and Yuanqing Xia and Bo Xiao",

note = "Publisher Copyright: {\textcopyright} 2023",

year = "2023",

month = sep,

doi = "10.1016/j.imavis.2023.104760",

language = "English",

volume = "137",

journal = "Image and Vision Computing",

issn = "0262-8856",

publisher = "Elsevier Ltd.",

}

TY - JOUR

T1 - Visual tracking using transformer with a combination of convolution and attention

AU - Wang, Yuxuan

AU - Yan, Liping

AU - Feng, Zihang

AU - Xia, Yuanqing

AU - Xiao, Bo

PY - 2023/9

Y1 - 2023/9

N2 - For Siamese-based trackers in the field of single object tracking, cross-correlation operation plays an important role. However, the cross-correlation essentially uses target feature to locally linearly match the search region, which leads to insufficient utilization or even loss of feature information. To effectively employ global context and sufficiently explore the relevance of template and search region, a novel matching operator is designed inspired by Transformer, which uses multi-head attention and embed a designed modulation module across the inputs of operator. Meanwhile, we equip our tracker with a multi-scale encoder/decoder strategy to gradually make more precise tracking. Finally, a complete tracking framework is presented named VTTR. The tracker consists of a feature extractor, a multi-scale encoder based on depth-wise convolution, a modified decoder as the matching operator and a prediction head. The proposed tracker is tested on many benchmarks and achieve excellent performance while running with fast speed.

AB - For Siamese-based trackers in the field of single object tracking, cross-correlation operation plays an important role. However, the cross-correlation essentially uses target feature to locally linearly match the search region, which leads to insufficient utilization or even loss of feature information. To effectively employ global context and sufficiently explore the relevance of template and search region, a novel matching operator is designed inspired by Transformer, which uses multi-head attention and embed a designed modulation module across the inputs of operator. Meanwhile, we equip our tracker with a multi-scale encoder/decoder strategy to gradually make more precise tracking. Finally, a complete tracking framework is presented named VTTR. The tracker consists of a feature extractor, a multi-scale encoder based on depth-wise convolution, a modified decoder as the matching operator and a prediction head. The proposed tracker is tested on many benchmarks and achieve excellent performance while running with fast speed.

KW - Attention

KW - Siamese networks

KW - Transformer

KW - Visual tracking

UR - http://www.scopus.com/inward/record.url?scp=85165055806&partnerID=8YFLogxK

U2 - 10.1016/j.imavis.2023.104760

DO - 10.1016/j.imavis.2023.104760

M3 - Article

AN - SCOPUS:85165055806

SN - 0262-8856

VL - 137

JO - Image and Vision Computing

JF - Image and Vision Computing

M1 - 104760

ER -

Visual tracking using transformer with a combination of convolution and attention

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this