TY - JOUR
T1 - Open-Vocabulary Object Detection via Cross-Modal Feature Alignment with CLIP and SAM
AU - Wang, Chongwen
AU - Xu, Hao
AU - Zheng, Zhiwei
N1 - Publisher Copyright:
© 2026, Beijing Institute of Technology. All rights reserved.
PY - 2026
Y1 - 2026
N2 - The growing demand for applications in open-world scenarios has driven progress in Open Vocabulary Object Detection (OVD) to overcome category constraints of traditional detection methods. While leveraging large models for region-text alignment has become a common OVD approach in recent years, it faces challenges such as limited category space, high computational cost, and domain gaps. In this paper, a cross-modal feature alignment method for OVD was proposed. Building on the CLIP and SAM models, semantic knowledge of CLIP was integrated with spatial perception of SAM through a bidirectional knowledge transfer architecture, extending “CLIP+SAM” from segmentation to detection. The cross-modal fusion module was plug-and-play and enabled rich multi-modal interaction during encoding and decoding, improving performance with minimal computational overhead. Experiments on COCO novel categories achieved an AP50 of 44.1, outperforming most existing methods and matching state-of-the-art results, while only increasing parameters by 10.3% without significantly raising training or inference costs.
AB - The growing demand for applications in open-world scenarios has driven progress in Open Vocabulary Object Detection (OVD) to overcome category constraints of traditional detection methods. While leveraging large models for region-text alignment has become a common OVD approach in recent years, it faces challenges such as limited category space, high computational cost, and domain gaps. In this paper, a cross-modal feature alignment method for OVD was proposed. Building on the CLIP and SAM models, semantic knowledge of CLIP was integrated with spatial perception of SAM through a bidirectional knowledge transfer architecture, extending “CLIP+SAM” from segmentation to detection. The cross-modal fusion module was plug-and-play and enabled rich multi-modal interaction during encoding and decoding, improving performance with minimal computational overhead. Experiments on COCO novel categories achieved an AP50 of 44.1, outperforming most existing methods and matching state-of-the-art results, while only increasing parameters by 10.3% without significantly raising training or inference costs.
KW - feature fusion
KW - object detection
KW - open vocabulary
KW - region-text alignment
UR - https://www.scopus.com/pages/publications/105041195305
U2 - 10.15918/j.tbit1001-0645.2025.159
DO - 10.15918/j.tbit1001-0645.2025.159
M3 - Article
AN - SCOPUS:105041195305
SN - 1001-0645
VL - 46
JO - Beijing Ligong Daxue Xuebao/Transaction of Beijing Institute of Technology
JF - Beijing Ligong Daxue Xuebao/Transaction of Beijing Institute of Technology
IS - 6
ER -