TY - GEN
T1 - Optimization of Cross-Domain Object Detection Based on RT-DETR
AU - Long, Zhiqi
AU - Chen, Wenjie
AU - Lin, Jiayi
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - To address the challenges of diverse object morphologies and significant inter-domain distribution differences in cross-domain object detection, this paper introduces the integration of Deformable Large Kernel Attention (DLKA) into the RT-DETR detection model, enhancing performance through the innovative combination of large convolution kernels and deformable convolutions. Specifically, large convolution kernels employ depth-wise separable convolutions and dilation techniques to expand the receptive field for capturing rich contextual information while reducing computational costs, effectively mimicking the global feature modeling capability of self-attention mechanisms. Deformable convolutions dynamically learn sampling offsets to adaptively adjust the sampling positions of convolution kernels, enhancing the model's adaptability to irregular object shapes and complex layouts in cross-domain scenarios. Experimental results demonstrate that the model incorporating DLKA achieves improvements of approximately 6.1%, 5.5%, and 5.6% in mAP50, Recall, and Precision metrics, respectively, compared to the baseline RT-DETR-resnet18. Notably, it exhibits more substantial performance advantages in late-stage training. This mechanism fundamentally enhances the model's robustness to object morphology and distribution differences across domains, significantly improving detection accuracy and recall in cross-domain scenarios. It provides an effective solution to address domain discrepancies in cross-domain object detection.
AB - To address the challenges of diverse object morphologies and significant inter-domain distribution differences in cross-domain object detection, this paper introduces the integration of Deformable Large Kernel Attention (DLKA) into the RT-DETR detection model, enhancing performance through the innovative combination of large convolution kernels and deformable convolutions. Specifically, large convolution kernels employ depth-wise separable convolutions and dilation techniques to expand the receptive field for capturing rich contextual information while reducing computational costs, effectively mimicking the global feature modeling capability of self-attention mechanisms. Deformable convolutions dynamically learn sampling offsets to adaptively adjust the sampling positions of convolution kernels, enhancing the model's adaptability to irregular object shapes and complex layouts in cross-domain scenarios. Experimental results demonstrate that the model incorporating DLKA achieves improvements of approximately 6.1%, 5.5%, and 5.6% in mAP50, Recall, and Precision metrics, respectively, compared to the baseline RT-DETR-resnet18. Notably, it exhibits more substantial performance advantages in late-stage training. This mechanism fundamentally enhances the model's robustness to object morphology and distribution differences across domains, significantly improving detection accuracy and recall in cross-domain scenarios. It provides an effective solution to address domain discrepancies in cross-domain object detection.
KW - Cross-domain Object Detection
KW - DLKA
KW - RT-DETR
KW - Transformer
UR - https://www.scopus.com/pages/publications/105040975891
U2 - 10.1109/CAC67268.2025.11487814
DO - 10.1109/CAC67268.2025.11487814
M3 - Conference contribution
AN - SCOPUS:105040975891
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 1057
EP - 1062
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 China Automation Congress, CAC 2025
Y2 - 26 September 2025 through 28 September 2025
ER -