TY - JOUR
T1 - ROARS
T2 - A Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis
AU - Zhang, Jingrui
AU - Liang, Feng
AU - Zhang, Yong
AU - Shangguan, Zixuan
AU - Li, Zhida
AU - Fan, Xiaoyi
AU - Leung, Victor C.M.
AU - Li, Guanbin
AU - Hu, Xiping
N1 - Publisher Copyright:
© 2013 IEEE.
PY - 2026
Y1 - 2026
N2 - The paradigm shift of remote sensing (RS) analysis toward cloud-based large-scale computing demands intelligent methods capable of interpreting petabyte-scale, high-resolution heterogeneous imagery. Multimodal Large Language Models (MLLMs) have emerged as a transformative solution for such data-intensive tasks, offering unprecedented capabilities in open-set perception and high-level semantic reasoning. However, directly applying general MLLM architectures to the RS domain faces significant challenges: the constraint of fixed-resolution vision encoders causes severe information loss for small objects, while the reliance on single-layer feature fusion neglects critical spatial cues preserved in intermediate layers. To address these challenges, we propose ROARS: a Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis, tailored for high-fidelity interpretation. First, we introduce a Native-Resolution Vision Encoder tailored to process images of arbitrary aspect ratios and scales without downsampling, thereby strictly preserving pixel-level details. Second, we design a Shortcut-Based Fusion Mechanism that bridges the semantic gap by injecting multi-level visual features into the language model. Furthermore, to enhance global contextual understanding, we incorporate a Multi-Granularity Token Pooling strategy that provides fine- and coarse-grained visual representations. Extensive experiments on the VRSBench demonstrate that our framework significantly outperforms state-of-the-art approaches, particularly in fine-grained object perception and complex spatial reasoning. These results validate the proposed method as a robust and efficient core component for next-generation cloud-based RS interpretation systems.
AB - The paradigm shift of remote sensing (RS) analysis toward cloud-based large-scale computing demands intelligent methods capable of interpreting petabyte-scale, high-resolution heterogeneous imagery. Multimodal Large Language Models (MLLMs) have emerged as a transformative solution for such data-intensive tasks, offering unprecedented capabilities in open-set perception and high-level semantic reasoning. However, directly applying general MLLM architectures to the RS domain faces significant challenges: the constraint of fixed-resolution vision encoders causes severe information loss for small objects, while the reliance on single-layer feature fusion neglects critical spatial cues preserved in intermediate layers. To address these challenges, we propose ROARS: a Cloud-Based Resolution-Native Framework for Large-Scale Remote Sensing Multimodal Analysis, tailored for high-fidelity interpretation. First, we introduce a Native-Resolution Vision Encoder tailored to process images of arbitrary aspect ratios and scales without downsampling, thereby strictly preserving pixel-level details. Second, we design a Shortcut-Based Fusion Mechanism that bridges the semantic gap by injecting multi-level visual features into the language model. Furthermore, to enhance global contextual understanding, we incorporate a Multi-Granularity Token Pooling strategy that provides fine- and coarse-grained visual representations. Extensive experiments on the VRSBench demonstrate that our framework significantly outperforms state-of-the-art approaches, particularly in fine-grained object perception and complex spatial reasoning. These results validate the proposed method as a robust and efficient core component for next-generation cloud-based RS interpretation systems.
KW - Multimodal Large Language Model
KW - Native Resolution
KW - Remote Sensing
KW - Vision-Language Information Fusion
UR - https://www.scopus.com/pages/publications/105046830554
U2 - 10.1109/TCC.2026.3721293
DO - 10.1109/TCC.2026.3721293
M3 - Article
AN - SCOPUS:105046830554
SN - 2168-7161
JO - IEEE Transactions on Cloud Computing
JF - IEEE Transactions on Cloud Computing
ER -