TY - JOUR
T1 - One size doesn’t fit all
T2 - Divide-and-conquer detector for UAV images
AU - Zhang, Yu
AU - Han, Yuqi
AU - Yang, Xiaojing
AU - Zhang, Xin
AU - Bao, Zengdi
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2027/1/1
Y1 - 2027/1/1
N2 - Object detection within unmanned aerial vehicle (UAV) visual perception systems grapples with core challenges including diverse detection difficulties across multi-scale targets, divergent optimization objectives between localization and classification tasks, and limited on-board computational resources. However, traditional detectors apply uniform model structures and capacities to multi-scale neck network and multi-task head network, resulting in unreasonable computational resource allocation that disrupts the balance between computation and performance. To address these issues, we propose a divide-and-conquer detector (DICDet). Firstly, a novel MRHNet as the backbone integrated with MRH module is proposed, which consists of multi-gradient flow, receptive field expansion, along with high-dimensional feature preservation. It includes two versions to enhance the network’s global perception and cross-channel correlation, effectively strengthening the representation ability of diverse targets in complex backgrounds. Secondly, a divide-and-conquer strategy guides the design of both the neck and head networks. Specifically, for the scale-specific neck network, structures of different computation are employed to process features of multi-size targets, properly allocating computational resources. For the task-specific head network, an asymmetric decoupled head with two specific task heads is constructed to meet the feature requirements for localization and classification tasks, respectively. Finally, we develop a new family of detectors with 5 model scales for UAV images: DICDet-N, S, M, L, and X. Experiments on the VisDrone2019-DET, AI-TOD-v2, and DOTA-v1.0 datasets prove that DICDet achieves higher accuracy with reduced computation, optimizing allocation of computational resources and achieving a comprehensive balance. The code is available at: https://github.com/PerSARption/DICDet.
AB - Object detection within unmanned aerial vehicle (UAV) visual perception systems grapples with core challenges including diverse detection difficulties across multi-scale targets, divergent optimization objectives between localization and classification tasks, and limited on-board computational resources. However, traditional detectors apply uniform model structures and capacities to multi-scale neck network and multi-task head network, resulting in unreasonable computational resource allocation that disrupts the balance between computation and performance. To address these issues, we propose a divide-and-conquer detector (DICDet). Firstly, a novel MRHNet as the backbone integrated with MRH module is proposed, which consists of multi-gradient flow, receptive field expansion, along with high-dimensional feature preservation. It includes two versions to enhance the network’s global perception and cross-channel correlation, effectively strengthening the representation ability of diverse targets in complex backgrounds. Secondly, a divide-and-conquer strategy guides the design of both the neck and head networks. Specifically, for the scale-specific neck network, structures of different computation are employed to process features of multi-size targets, properly allocating computational resources. For the task-specific head network, an asymmetric decoupled head with two specific task heads is constructed to meet the feature requirements for localization and classification tasks, respectively. Finally, we develop a new family of detectors with 5 model scales for UAV images: DICDet-N, S, M, L, and X. Experiments on the VisDrone2019-DET, AI-TOD-v2, and DOTA-v1.0 datasets prove that DICDet achieves higher accuracy with reduced computation, optimizing allocation of computational resources and achieving a comprehensive balance. The code is available at: https://github.com/PerSARption/DICDet.
KW - Aerial images
KW - Decoupled head
KW - Divide-and-conquer
KW - Multi-scale targets
KW - Object detection
UR - https://www.scopus.com/pages/publications/105046253275
U2 - 10.1016/j.eswa.2026.133761
DO - 10.1016/j.eswa.2026.133761
M3 - Article
AN - SCOPUS:105046253275
SN - 0957-4174
VL - 333
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 133761
ER -