Skip to main navigation Skip to search Skip to main content

Enhanced Head: Exploring Strong Detection Heads With Vision Transformer

  • Zewen Du
  • , Zhenjiang Hu
  • , Guiyu Zhao
  • , Ying Jin
  • , Hongbin Ma*
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

As a crucial component of object detectors, current detection heads often lack the capability to effectively utilize contextual information, adapt to deformable objects, and align features and tasks. However, most existing methods prioritize a single capability, lacking comprehensive approaches to introduce them simultaneously. In this paper, we propose the Enhanced Head to integrate the above three capabilities into the detectors concurrently. Specifically, we propose three attention blocks with linear complexity: Global Concentrated Attention (GCA), Local Deformable Cross-Task Attention (LDCA), and Boundary-Aware Cross-Task Attention (BACA). The GCA captures long-range dependencies efficiently by employing Spatial Information Concentration (SIC). The LDCA improves feature alignment and deformation adaptability by enabling local deformable cross-task feature interactions. The BACA aligns classification features with localization results, enhancing task alignment and further improving deformation adaptability through a region-deformable interaction scheme. We implement Enhanced Head as a plug-and-play detection head and evaluate its effectiveness through extensive experiments on the MS COCO and VisDrone datasets. For instance, on the COCO detection benchmark, our Enhanced Head achieves +3.6 AP gain for FSAF, +3.3 AP for RetinaNet, and +2.9 AP for ATSS while reducing the FLOPs.

Original languageEnglish
Pages (from-to)7834-7848
Number of pages15
JournalIEEE Transactions on Multimedia
Volume27
DOIs
Publication statusPublished - 2025
Externally publishedYes

Keywords

  • Detection head
  • object detection
  • vision transformer

Fingerprint

Dive into the research topics of 'Enhanced Head: Exploring Strong Detection Heads With Vision Transformer'. Together they form a unique fingerprint.

Cite this