Skip to main navigation Skip to search Skip to main content

RAST: Reliability-Aware Spatiotemporal Modeling for Multiobject Tracking in Satellite Videos

  • Yuting Shi
  • , Guanqun Wang*
  • , Yin Zhuang
  • , Tong Zhang
  • , He Chen
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Peking University

Research output: Contribution to journalArticlepeer-review

Abstract

Satellite remote sensing videos enable persistent wide-area observation, yet multiobject tracking (MOT) in this domain remains challenging due to tiny targets, cluttered backgrounds, subtle and nonlinear displacements, and time-varying observation quality. Such reliability-variant conditions can induce temporal representation drift, intensify multitask interference between detection and motion estimation, and make conventional reliability-invariant association brittle, leading to fragmented trajectories and identity switches. From a unified reliability-aware spatiotemporal modeling perspective, we propose RAST, an online and causal tracking framework that progressively improves robustness from representation learning to data association. RAST consists of three complementary components: 1) a dual-view temporal enhancement mechanism (DVTEM) that couples forward inertia propagation with backward contextual verification within a causal window to suppress drift accumulation; 2) a motion-guided dynamic feature aggregation mechanism (MGDFAM) that constructs motion priors from a temporal feature bank and performs task-specific deformable attention with decoupled representations for detection and displacement estimation; and 3) a hybrid-driven cascade association strategy (HDCAS) that leverages detection confidence as a proxy of observation reliability to adaptively switch between displacement-guided matching and Kalman-filter-based inertial association. Extensive experiments on two satellite video MOT benchmarks, SatVideoDT and SatMTB-MOT, demonstrate that RAST consistently outperforms state-of-the-art (SOTA) methods. In particular, RAST achieves 67.0% MOTA and 75.1% Rcll on SatVideoDT, and attains the best MOTA across all three subsets of SatMTB-MOT, validating its robustness under crowded tiny-object scenarios and multiclass variations.

Original languageEnglish
Article number5628419
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume64
DOIs
Publication statusPublished - 2026
Externally publishedYes

Keywords

  • Multiobject tracking (MOT)
  • reliability-aware
  • satellite video
  • spatiotemporal modeling

Fingerprint

Dive into the research topics of 'RAST: Reliability-Aware Spatiotemporal Modeling for Multiobject Tracking in Satellite Videos'. Together they form a unique fingerprint.

Cite this