TY - GEN
T1 - Optimization of Cross-Domain Detection Capabilities Based on RT-DETR
AU - Cao, Yue
AU - Chen, Wenjie
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Object detectors often suffer a significant performance decline when faced with domain shifts between the source domain (collected data) and the target domain (actual application data). This is due to significant visual differences between images across domains, such as variations in object scale, texture, and content style. To improve cross-domain detection performance, this paper proposes integrating two modules: AssemFormer (an assembly-based convolutional vision transformer) and SEAM (Separated and Enhanced Attention Module) into the RT-DETR detector. AssemFormer combines the local feature extraction capabilities of convolutional neural networks with the global context modeling power of Transformers, addressing the limitations of traditional convolutional neural networks in capturing long-range dependencies and local details. SEAM improves feature responses in unobstructed regions while compensating for information loss in occluded areas, thereby enhancing detection capabilities for obscured objects. It also addresses the lack of inductive bias and weak local detail capture in pure Transformers. Together, these modules mitigate the adverse effects of domain differences between synthetic and real images, optimizing performance for cross-domain object detection. In the Sim10k-Cityscapes crossdomain detection task, the mAP improved by 5.8%, and in the Cityscapes-FoggyCityscapes task, it increased by 5.7%.
AB - Object detectors often suffer a significant performance decline when faced with domain shifts between the source domain (collected data) and the target domain (actual application data). This is due to significant visual differences between images across domains, such as variations in object scale, texture, and content style. To improve cross-domain detection performance, this paper proposes integrating two modules: AssemFormer (an assembly-based convolutional vision transformer) and SEAM (Separated and Enhanced Attention Module) into the RT-DETR detector. AssemFormer combines the local feature extraction capabilities of convolutional neural networks with the global context modeling power of Transformers, addressing the limitations of traditional convolutional neural networks in capturing long-range dependencies and local details. SEAM improves feature responses in unobstructed regions while compensating for information loss in occluded areas, thereby enhancing detection capabilities for obscured objects. It also addresses the lack of inductive bias and weak local detail capture in pure Transformers. Together, these modules mitigate the adverse effects of domain differences between synthetic and real images, optimizing performance for cross-domain object detection. In the Sim10k-Cityscapes crossdomain detection task, the mAP improved by 5.8%, and in the Cityscapes-FoggyCityscapes task, it increased by 5.7%.
KW - Algorithm Optimization
KW - Feature Extraction
KW - Object Detection
KW - RT-DETR
UR - https://www.scopus.com/pages/publications/105047336558
U2 - 10.1109/ICCA69928.2026.11617994
DO - 10.1109/ICCA69928.2026.11617994
M3 - Conference contribution
AN - SCOPUS:105047336558
T3 - IEEE International Conference on Control and Automation, ICCA
SP - 174
EP - 179
BT - 2026 IEEE 20th International Conference on Control and Automation, ICCA 2026
PB - IEEE Computer Society
T2 - 20th IEEE International Conference on Control and Automation, ICCA 2026
Y2 - 16 June 2026 through 19 June 2026
ER -