TY - JOUR
T1 - CMD3
T2 - Cross-Modal Decoupled Deformable Distillation for EEG-fNIRS Fusion
AU - Fan, Tianqi
AU - Tian, Fuze
AU - Wang, Su
AU - Zhang, Haoyan
AU - Luo, Gang
AU - Zhu, Lixian
AU - Liu, Jingxin
AU - Jiang, Lei
AU - Cai, Ran
AU - Dong, Qunxi
AU - Hu, Bin
N1 - Publisher Copyright:
© 2010-2012 IEEE.
PY - 2026/4
Y1 - 2026/4
N2 - Multimodal fusion of Electroencephalography (EEG) and functional Near-Infrared Spectroscopy (fNIRS) has shown great promise in Brain-Computer Interface (BCI) tasks. However, due to differences in physical mechanisms, temporal dynamics, and semantic representations between the two modalities, the fusion process faces significant challenges such as heterogeneity and temporal misalignment. To address this, we propose a cross-modal decoupled deformable distillation (CMD3) method, which aims to achieve flexible, efficient, and interpretable EEG-fNIRS fusion learning. CMD3 first decouples the feature representations of each modality into modality-independent and modality-specific spaces to separately model commonality and complementary information. A deformable feature extraction network is then designed to process shared and specific features individually, enabling cross-modal temporal alignment via predicted dynamic offsets, thereby mitigating response delays between modalities. Furthermore, to facilitate inter-modal knowledge transfer, we construct a dual-space graph distillation module to explicitly migrate semantic information across modalities, with learnable edge weights used to adaptively regulate the distillation strength. CMD3 is systematically evaluated on public datasets covering emotion recognition and motor imagery tasks. Experimental results demonstrate that CMD3 consistently outperforms existing fusion approaches in classification performance. Offset visualization further reveals physiologically meaningful temporal attention patterns learned by the model, validating the effectiveness and explainability of the proposed method.
AB - Multimodal fusion of Electroencephalography (EEG) and functional Near-Infrared Spectroscopy (fNIRS) has shown great promise in Brain-Computer Interface (BCI) tasks. However, due to differences in physical mechanisms, temporal dynamics, and semantic representations between the two modalities, the fusion process faces significant challenges such as heterogeneity and temporal misalignment. To address this, we propose a cross-modal decoupled deformable distillation (CMD3) method, which aims to achieve flexible, efficient, and interpretable EEG-fNIRS fusion learning. CMD3 first decouples the feature representations of each modality into modality-independent and modality-specific spaces to separately model commonality and complementary information. A deformable feature extraction network is then designed to process shared and specific features individually, enabling cross-modal temporal alignment via predicted dynamic offsets, thereby mitigating response delays between modalities. Furthermore, to facilitate inter-modal knowledge transfer, we construct a dual-space graph distillation module to explicitly migrate semantic information across modalities, with learnable edge weights used to adaptively regulate the distillation strength. CMD3 is systematically evaluated on public datasets covering emotion recognition and motor imagery tasks. Experimental results demonstrate that CMD3 consistently outperforms existing fusion approaches in classification performance. Offset visualization further reveals physiologically meaningful temporal attention patterns learned by the model, validating the effectiveness and explainability of the proposed method.
KW - EEG
KW - cross-modal
KW - deformable alignment
KW - fNIRS
KW - graph distillation
UR - https://www.scopus.com/pages/publications/105032488817
U2 - 10.1109/TAFFC.2026.3670503
DO - 10.1109/TAFFC.2026.3670503
M3 - Article
AN - SCOPUS:105032488817
SN - 1949-3045
VL - 17
SP - 2012
EP - 2026
JO - IEEE Transactions on Affective Computing
JF - IEEE Transactions on Affective Computing
IS - 2
ER -