Skip to main navigation Skip to search Skip to main content

Modulating Dense Content with Sparse Context for Real-time Visual Recognition

  • Beijing Institute of Technology
  • Hunan University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Efficient vision backbones face a fundamental conflict between expanding the receptive field to enhance semantic understanding and reducing the computational cost of real-time inference. While Vision Transformers (ViTs) and recently revived Large-Kernel CNNs can capture global contextual information, they often suffer from excessive redundancy and feature over-smoothing. To resolve this, we propose the Asymmetric Contextual Modulation Convolution (ACMConv), a novel operator that introduces an asymmetric dual-branch design. The Context Modulator branch utilizes multi-scale dilated convolutions to aggregate long-range contextual information with minimal cost. The Content Descriptor branch employs dense convolutions to strictly preserve local details. By modulating the dense content with the sparse context, we effectively mitigate the detail loss common in large-kernel paradigms. Based on this core operator, we construct ACMNet, a pure ConvNet architecture that prioritizes role-specific efficiency. During inference, this multi-branch structure can be reparameterized into standard single convolutions, ensuring contiguous memory access and high throughput. ACMNet achieves state-of-the-art performance on ImageNet-1K, COCO object detection, and instance segmentation. Furthermore, our ACMNet-T model achieves a latency of 8.9 ms on an Intel i7-12700KF CPU, demonstrating real-time efficiency on commodity hardware.

Original languageEnglish
Title of host publication2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798319531193
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026 - Xuzhou, China
Duration: 8 May 202610 May 2026

Publication series

Name2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026

Conference

Conference2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
Country/TerritoryChina
CityXuzhou
Period8/05/2610/05/26

Keywords

  • Efficient Vision Backbone
  • Large-Kernel CNN
  • Real-time Visual Recognition
  • Structural Reparameterization

Fingerprint

Dive into the research topics of 'Modulating Dense Content with Sparse Context for Real-time Visual Recognition'. Together they form a unique fingerprint.

Cite this