TY - JOUR
T1 - LiteMFT
T2 - Lightweight Multi-Modal Fine-Tuning for Semantic Segmentation
AU - Guo, Chengwang
AU - Zhang, Yuxiang
AU - Zhang, Mengmeng
AU - Liu, Huan
AU - Li, Wei
N1 - Publisher Copyright:
© 2026 IEEE. All rights reserved.
PY - 2026
Y1 - 2026
N2 - Multi-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone’s adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multimodal semantic segmentation.
AB - Multi-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone’s adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multimodal semantic segmentation.
KW - lightweight fine-tuning
KW - Multi-modal semantic segmentation
KW - vision foundation models
UR - https://www.scopus.com/pages/publications/105043085801
U2 - 10.1109/TIP.2026.3702423
DO - 10.1109/TIP.2026.3702423
M3 - Article
C2 - 42313582
AN - SCOPUS:105043085801
SN - 1057-7149
VL - 35
SP - 6674
EP - 6685
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
ER -