Skip to main navigation Skip to search Skip to main content

MCRVT: Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition

  • Xin Heng Li
  • , Zhen Tao Liu*
  • , Yu Jie Zou
  • , Jinhua She
  • , Kaoru Hirota
  • *Corresponding author for this work
  • China University of Geosciences, Wuhan
  • Tokyo University of Technology
  • Institute of Science Tokyo

Research output: Contribution to journalArticlepeer-review

Abstract

Two challenges for Speech Emotion Recognition (SER) are efficiently capturing features to explore speech-emotion correlations and reducing the SER model redundancy while increasing performance. To address these challenges, we propose Multi-Hierarchical Cross-Reconstruction Networks with Versatile Transformer (MCRVT), including a Spectrogram-Based SER Model with Versatile Attention (SSER-VA), a pre-trained model (WavLM), and Multi-Hierarchical Feature Fusion with Cross Attention (MHFF-CA). Specifically, SSER-VA takes the log-Mel spectrogram as the input, and we propose a versatile attention module to reduce redundancy and increase efficiency. MHFF-CA is employed to refine intermediate features from WavLM and SSER-VA to improve competitiveness in various tasks, where Contrastive Reconstruction Networks (CRN) and Cross Attention-based Feature Fusion (CAF) are utilized to reconstruct and fuse features, respectively. Effectiveness experiments on IEMOCAP and MELD datasets illustrate MCRVT's superior performance compared to SOTA methods. Exploratory experiments on D-vlog, BioCAS2024, and IEMOCAP (speaker-independent setting) datasets demonstrate the competitiveness of our method in various speech classification tasks.

Original languageEnglish
Pages (from-to)2189-2199
Number of pages11
JournalIEEE Transactions on Affective Computing
Volume16
Issue number3
DOIs
Publication statusPublished - 2025
Externally publishedYes

Keywords

  • Speech emotion recognition
  • cross-reconstruction networks
  • speech depression recognition
  • transformer
  • versatile attention

Fingerprint

Dive into the research topics of 'MCRVT: Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition'. Together they form a unique fingerprint.

Cite this