Skip to main navigation Skip to search Skip to main content

ROARS: A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis

  • Jingrui Zhang
  • , Feng Liang*
  • , Yong Zhang
  • , Zixuan Shangguan
  • , Zhida Li
  • , Xiaoyi Fan
  • , Victor C.M. Leung
  • , Guanbin Li
  • , Xiping Hu
  • *Corresponding author for this work
  • Shenzhen MSU-BIT University
  • Beijing Institute of Technology
  • Sun Yat-Sen University
  • Guangdong Province Key Laboratory of Big Data Analysis and Processing
  • University of British Columbia

Research output: Contribution to journalArticlepeer-review

Abstract

The paradigm shift of remote sensing (RS) analysis toward cloud-based large-scale computing demands intelligent methods capable of interpreting petabyte-scale, high-resolution heterogeneous imagery. Multimodal Large Language Models (MLLMs) have emerged as a transformative solution for such data-intensive tasks, offering unprecedented capabilities in open-set perception and high-level semantic reasoning. However, directly applying general MLLM architectures to the RS domain faces significant challenges: the constraint of fixed-resolution vision encoders causes severe information loss for small objects, while the reliance on single-layer feature fusion neglects critical spatial cues preserved in intermediate layers. To address these challenges, we propose ROARS: a Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis, tailored for high-fidelity interpretation. First, we introduce a Native-Resolution Vision Encoder tailored to process images of arbitrary aspect ratios and scales without downsampling, thereby strictly preserving pixel-level details. Second, we design a Shortcut-Based Fusion Mechanism that bridges the semantic gap by injecting multi-level visual features into the language model. Furthermore, to enhance global contextual understanding, we incorporate a Multi-Granularity Token Pooling strategy that provides fine- and coarse-grained visual representations. Extensive experiments on the VRSBench demonstrate that our framework significantly outperforms state-of-the-art approaches, particularly in fine-grained object perception and complex spatial reasoning. These results validate the proposed method as a robust and efficient core component for next-generation cloud-based RS interpretation systems.

Original languageEnglish
JournalIEEE Transactions on Cloud Computing
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Multimodal Large Language Model
  • Native Resolution
  • Remote Sensing
  • Vision-Language Information Fusion

Fingerprint

Dive into the research topics of 'ROARS: A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis'. Together they form a unique fingerprint.

Cite this