Controllable Generative Knowledge-Driven Few-Shot Object Detection from Optical Remote Sensing Imagery

Tong Zhang, Yin Zhuang*, Guanqun Wang, He Chen, Hao Wang, Lianlin Li, Jun Li

*此作品的通讯作者

科研成果: 期刊稿件文章同行评审

摘要

Few-shot object detection (FSOD) has to learn classification and localization information for unseen object detection under very low-data resource regimes. However, when deficient samples are adopted for model training, it is hard to build powerful location-Aware and identification abilities for well coping with agnostic bias from diverse testing scenarios; at the same time, the overfitting phenomenon is easily occurring. Therefore, in this article, a controllable generative knowledge-driven FSOD called CGK-FSOD is proposed for unseen object detection from optical remote sensing imagery. Specifically, to enrich the learnable data space of scarce samples for preventing incomplete agnostic-bias learning, while avoiding the overfitting phenomenon, a visual-Textual prompt-based controllable data generation is designed to generate high-quality object detection data based on pretrained foundational models [i.e., the stable diffusion (SD) and contrastive language-image pre-Training (CLIP)], which not only can introduce the generalized domain-level knowledge into the remote sensing domain but also sets up an all-round data space to support complete learning of potential agnostic bias. Furthermore, with respect to the denoising generative process of SD, a series of cross-modality generative features in latent representation space are reused for few-shot fine-Tuning by the designed cross-modality feature embedding (CMFE), which not only can bring diverse generative abilities into the feature fusion step of the detector but also gracefully sets up feature representation scalability to make the detector better adapt to agnostic bias from diverse testing scenarios of FSOD. Finally, extensive experiments are executed on two public remote sensing datasets (e.g., DIOR and NWPUVHR-10), and the results indicate that the proposed CGK-FSOD is very effective and flexible for FSOD.

源语言英语
文章编号5612319
期刊IEEE Transactions on Geoscience and Remote Sensing
63
DOI
出版状态已出版 - 2025

指纹

探究 'Controllable Generative Knowledge-Driven Few-Shot Object Detection from Optical Remote Sensing Imagery' 的科研主题。它们共同构成独一无二的指纹。

引用此

Zhang, T., Zhuang, Y., Wang, G., Chen, H., Wang, H., Li, L., & Li, J. (2025). Controllable Generative Knowledge-Driven Few-Shot Object Detection from Optical Remote Sensing Imagery. IEEE Transactions on Geoscience and Remote Sensing, 63, 文章 5612319. https://doi.org/10.1109/TGRS.2025.3541937