TY - GEN
T1 - Modulating Dense Content with Sparse Context for Real-time Visual Recognition
AU - Wang, Xuyang
AU - Miao, Lingjuan
AU - Zi, Yutian
AU - Zhou, Zhiqiang
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Efficient vision backbones face a fundamental conflict between expanding the receptive field to enhance semantic understanding and reducing the computational cost of real-time inference. While Vision Transformers (ViTs) and recently revived Large-Kernel CNNs can capture global contextual information, they often suffer from excessive redundancy and feature over-smoothing. To resolve this, we propose the Asymmetric Contextual Modulation Convolution (ACMConv), a novel operator that introduces an asymmetric dual-branch design. The Context Modulator branch utilizes multi-scale dilated convolutions to aggregate long-range contextual information with minimal cost. The Content Descriptor branch employs dense convolutions to strictly preserve local details. By modulating the dense content with the sparse context, we effectively mitigate the detail loss common in large-kernel paradigms. Based on this core operator, we construct ACMNet, a pure ConvNet architecture that prioritizes role-specific efficiency. During inference, this multi-branch structure can be reparameterized into standard single convolutions, ensuring contiguous memory access and high throughput. ACMNet achieves state-of-the-art performance on ImageNet-1K, COCO object detection, and instance segmentation. Furthermore, our ACMNet-T model achieves a latency of 8.9 ms on an Intel i7-12700KF CPU, demonstrating real-time efficiency on commodity hardware.
AB - Efficient vision backbones face a fundamental conflict between expanding the receptive field to enhance semantic understanding and reducing the computational cost of real-time inference. While Vision Transformers (ViTs) and recently revived Large-Kernel CNNs can capture global contextual information, they often suffer from excessive redundancy and feature over-smoothing. To resolve this, we propose the Asymmetric Contextual Modulation Convolution (ACMConv), a novel operator that introduces an asymmetric dual-branch design. The Context Modulator branch utilizes multi-scale dilated convolutions to aggregate long-range contextual information with minimal cost. The Content Descriptor branch employs dense convolutions to strictly preserve local details. By modulating the dense content with the sparse context, we effectively mitigate the detail loss common in large-kernel paradigms. Based on this core operator, we construct ACMNet, a pure ConvNet architecture that prioritizes role-specific efficiency. During inference, this multi-branch structure can be reparameterized into standard single convolutions, ensuring contiguous memory access and high throughput. ACMNet achieves state-of-the-art performance on ImageNet-1K, COCO object detection, and instance segmentation. Furthermore, our ACMNet-T model achieves a latency of 8.9 ms on an Intel i7-12700KF CPU, demonstrating real-time efficiency on commodity hardware.
KW - Efficient Vision Backbone
KW - Large-Kernel CNN
KW - Real-time Visual Recognition
KW - Structural Reparameterization
UR - https://www.scopus.com/pages/publications/105044249665
U2 - 10.1109/ICAISISAS68969.2026.11567783
DO - 10.1109/ICAISISAS68969.2026.11567783
M3 - Conference contribution
AN - SCOPUS:105044249665
T3 - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
BT - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
Y2 - 8 May 2026 through 10 May 2026
ER -