Skip to main navigation Skip to search Skip to main content

LiteMFT: Lightweight Multi-Modal Fine-Tuning for Semantic Segmentation

  • Beijing Institute of Technology
  • The University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-modal image segmentation has recently attracted considerable attention due to its ability to integrate complementary information from diverse sensors, thereby enabling more accurate semantic predictions in complex or specialized scenarios. However, as data volume and model capacity continue to grow, many existing methods suffer substantial increases in parameters and computational costs, particularly with the widespread adoption of Vision Foundation Models (VFMs). To address these challenges, we introduce a Lightweight Multi-modal Fine-Tuning framework (LiteMFT) designed for efficient and generalizable adaptation of RGB-pretrained VFMs to multi-modal semantic segmentation. By incorporating only a small number of trainable parameters, LiteMFT enables effective extension of existing models to handle multi-modal image fusion tasks. The framework centers around two key components: the Modality Local Competition (MLC) module, which dynamically and efficiently fuses complementary features across modalities, and the Gated Low-Rank Adapter (GLR), which improves the backbone’s adaptability to multi-modal data through content-aware low-rank transformation. Extensive experiments on both bi-modal and tri-modal segmentation tasks demonstrate that LiteMFT not only achieves competitive or superior performance but also exhibits strong scalability for additional modalities, underscoring its practicality and broad applicability in multimodal semantic segmentation.

Original languageEnglish
Pages (from-to)6674-6685
Number of pages12
JournalIEEE Transactions on Image Processing
Volume35
DOIs
Publication statusPublished - 2026

Keywords

  • lightweight fine-tuning
  • Multi-modal semantic segmentation
  • vision foundation models

Fingerprint

Dive into the research topics of 'LiteMFT: Lightweight Multi-Modal Fine-Tuning for Semantic Segmentation'. Together they form a unique fingerprint.

Cite this