TY - JOUR
T1 - MCRVT
T2 - Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition
AU - Li, Xin Heng
AU - Liu, Zhen Tao
AU - Zou, Yu Jie
AU - She, Jinhua
AU - Hirota, Kaoru
N1 - Publisher Copyright:
© 2010-2012 IEEE.
PY - 2025
Y1 - 2025
N2 - Two challenges for Speech Emotion Recognition (SER) are efficiently capturing features to explore speech-emotion correlations and reducing the SER model redundancy while increasing performance. To address these challenges, we propose Multi-Hierarchical Cross-Reconstruction Networks with Versatile Transformer (MCRVT), including a Spectrogram-Based SER Model with Versatile Attention (SSER-VA), a pre-trained model (WavLM), and Multi-Hierarchical Feature Fusion with Cross Attention (MHFF-CA). Specifically, SSER-VA takes the log-Mel spectrogram as the input, and we propose a versatile attention module to reduce redundancy and increase efficiency. MHFF-CA is employed to refine intermediate features from WavLM and SSER-VA to improve competitiveness in various tasks, where Contrastive Reconstruction Networks (CRN) and Cross Attention-based Feature Fusion (CAF) are utilized to reconstruct and fuse features, respectively. Effectiveness experiments on IEMOCAP and MELD datasets illustrate MCRVT's superior performance compared to SOTA methods. Exploratory experiments on D-vlog, BioCAS2024, and IEMOCAP (speaker-independent setting) datasets demonstrate the competitiveness of our method in various speech classification tasks.
AB - Two challenges for Speech Emotion Recognition (SER) are efficiently capturing features to explore speech-emotion correlations and reducing the SER model redundancy while increasing performance. To address these challenges, we propose Multi-Hierarchical Cross-Reconstruction Networks with Versatile Transformer (MCRVT), including a Spectrogram-Based SER Model with Versatile Attention (SSER-VA), a pre-trained model (WavLM), and Multi-Hierarchical Feature Fusion with Cross Attention (MHFF-CA). Specifically, SSER-VA takes the log-Mel spectrogram as the input, and we propose a versatile attention module to reduce redundancy and increase efficiency. MHFF-CA is employed to refine intermediate features from WavLM and SSER-VA to improve competitiveness in various tasks, where Contrastive Reconstruction Networks (CRN) and Cross Attention-based Feature Fusion (CAF) are utilized to reconstruct and fuse features, respectively. Effectiveness experiments on IEMOCAP and MELD datasets illustrate MCRVT's superior performance compared to SOTA methods. Exploratory experiments on D-vlog, BioCAS2024, and IEMOCAP (speaker-independent setting) datasets demonstrate the competitiveness of our method in various speech classification tasks.
KW - Speech emotion recognition
KW - cross-reconstruction networks
KW - speech depression recognition
KW - transformer
KW - versatile attention
UR - https://www.scopus.com/pages/publications/105002397188
U2 - 10.1109/TAFFC.2025.3557153
DO - 10.1109/TAFFC.2025.3557153
M3 - Article
AN - SCOPUS:105002397188
SN - 1949-3045
VL - 16
SP - 2189
EP - 2199
JO - IEEE Transactions on Affective Computing
JF - IEEE Transactions on Affective Computing
IS - 3
ER -