Skip to main navigation Skip to search Skip to main content

DCP-Net: Learning Detail–Context Perception via Spatial-Frequency Guidance for Tiny Object Detection in Remote Sensing Images

  • Beijing Institute of Technology
  • National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing
  • Beijing Institute of Remote Sensing Information

Research output: Contribution to journalArticlepeer-review

Abstract

Tiny object detection in remote sensing images haslong been challenging due to the low resolution of tiny objectsand complex backgrounds. However, existing adaptive receptivefield methods mainly focus on object and contextual information,making them susceptible to interference from complex backgrounds. This interference exacerbates the imbalance in sampleassignment and hinders accurate adaptation of receptive fieldsto object scales, resulting in insufficient representation of finegrained object details. In addition, most loss functions are proneto gradient instability in scenarios involving small objects. Toaddress the aforementioned problems, this paper proposes aDCP-Net, which learns detail–context perception under spatialfrequency guidance. To enhance the representation of detailsof tiny objects, the network is the first to leverage spatialfrequency to distinguish the structural features of objects andbackground. By introducing a spatial-frequency guided dualdilated convolution, it effectively enhances local object detailsand global semantic associations while suppressing interferencefrom complex backgrounds. Based on this convolutional unit, wefurther design a cross-layer fine-grained fusion module to recoverlost object details at a lower computational cost. Furthermore, weintroduce a Gaussian classification-regression loss to address theissues of imbalanced sample allocation and unstable gradient,while applying a decoupled probabilistic metric strategy anda multi-scale enhanced decoupled head to improve localizationaccuracy. Extensive experimental results demonstrate the effectiveness of the proposed method. Specifically, DCP-Net achievesan average precision (AP) of 27.5% on the AI-TODv2 datasetand an AP50 of 78.0% on the DOTA-v1.0 dataset.

Original languageEnglish
JournalIEEE Transactions on Geoscience and Remote Sensing
DOIs
Publication statusAccepted/In press - 2026

Keywords

  • multi-scale fusion
  • remote sensing images
  • Spatial-frequency guidance
  • tiny object detection (TOD)

Fingerprint

Dive into the research topics of 'DCP-Net: Learning Detail–Context Perception via Spatial-Frequency Guidance for Tiny Object Detection in Remote Sensing Images'. Together they form a unique fingerprint.

Cite this