跳到主要导航 跳到搜索 跳到主要内容

ROARS: A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis

  • Jingrui Zhang
  • , Feng Liang*
  • , Yong Zhang
  • , Zixuan Shangguan
  • , Zhida Li
  • , Xiaoyi Fan
  • , Victor C.M. Leung
  • , Guanbin Li
  • , Xiping Hu
  • *此作品的通讯作者
  • Shenzhen MSU-BIT University
  • Beijing Institute of Technology
  • Sun Yat-Sen University
  • Guangdong Province Key Laboratory of Big Data Analysis and Processing
  • University of British Columbia

科研成果: 期刊稿件文章同行评审

摘要

The paradigm shift of remote sensing (RS) analysis toward cloud-based large-scale computing demands intelligent methods capable of interpreting petabyte-scale, high-resolution heterogeneous imagery. Multimodal Large Language Models (MLLMs) have emerged as a transformative solution for such data-intensive tasks, offering unprecedented capabilities in open-set perception and high-level semantic reasoning. However, directly applying general MLLM architectures to the RS domain faces significant challenges: the constraint of fixed-resolution vision encoders causes severe information loss for small objects, while the reliance on single-layer feature fusion neglects critical spatial cues preserved in intermediate layers. To address these challenges, we propose ROARS: a Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis, tailored for high-fidelity interpretation. First, we introduce a Native-Resolution Vision Encoder tailored to process images of arbitrary aspect ratios and scales without downsampling, thereby strictly preserving pixel-level details. Second, we design a Shortcut-Based Fusion Mechanism that bridges the semantic gap by injecting multi-level visual features into the language model. Furthermore, to enhance global contextual understanding, we incorporate a Multi-Granularity Token Pooling strategy that provides fine- and coarse-grained visual representations. Extensive experiments on the VRSBench demonstrate that our framework significantly outperforms state-of-the-art approaches, particularly in fine-grained object perception and complex spatial reasoning. These results validate the proposed method as a robust and efficient core component for next-generation cloud-based RS interpretation systems.

源语言英语
期刊IEEE Transactions on Cloud Computing
DOI
出版状态已接受/待刊 - 2026
已对外发布

学术指纹

探究 'ROARS: A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis' 的科研主题。它们共同构成独一无二的学术指纹。

引用此