跳到主要导航 跳到搜索 跳到主要内容

Open-Vocabulary Object Detection via Cross-Modal Feature Alignment with CLIP and SAM

投稿的翻译标题: 基于CLIP-SAM的跨模态特征对齐开放词汇目标检测
  • Beijing Institute of Technology
  • China University of Labor Relations

科研成果: 期刊稿件文章同行评审

摘要

The growing demand for applications in open-world scenarios has driven progress in Open Vocabulary Object Detection (OVD) to overcome category constraints of traditional detection methods. While leveraging large models for region-text alignment has become a common OVD approach in recent years, it faces challenges such as limited category space, high computational cost, and domain gaps. In this paper, a cross-modal feature alignment method for OVD was proposed. Building on the CLIP and SAM models, semantic knowledge of CLIP was integrated with spatial perception of SAM through a bidirectional knowledge transfer architecture, extending “CLIP+SAM” from segmentation to detection. The cross-modal fusion module was plug-and-play and enabled rich multi-modal interaction during encoding and decoding, improving performance with minimal computational overhead. Experiments on COCO novel categories achieved an AP50 of 44.1, outperforming most existing methods and matching state-of-the-art results, while only increasing parameters by 10.3% without significantly raising training or inference costs.

投稿的翻译标题基于CLIP-SAM的跨模态特征对齐开放词汇目标检测
源语言英语
期刊Beijing Ligong Daxue Xuebao/Transaction of Beijing Institute of Technology
46
6
DOI
出版状态已出版 - 2026
已对外发布

指纹

探究 '基于CLIP-SAM的跨模态特征对齐开放词汇目标检测' 的科研主题。它们共同构成独一无二的指纹。

引用此