跳到主要导航 跳到搜索 跳到主要内容

MCRVT: Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition

  • Xin Heng Li
  • , Zhen Tao Liu*
  • , Yu Jie Zou
  • , Jinhua She
  • , Kaoru Hirota
  • *此作品的通讯作者
  • China University of Geosciences, Wuhan
  • Tokyo University of Technology
  • Institute of Science Tokyo

科研成果: 期刊稿件文章同行评审

摘要

Two challenges for Speech Emotion Recognition (SER) are efficiently capturing features to explore speech-emotion correlations and reducing the SER model redundancy while increasing performance. To address these challenges, we propose Multi-Hierarchical Cross-Reconstruction Networks with Versatile Transformer (MCRVT), including a Spectrogram-Based SER Model with Versatile Attention (SSER-VA), a pre-trained model (WavLM), and Multi-Hierarchical Feature Fusion with Cross Attention (MHFF-CA). Specifically, SSER-VA takes the log-Mel spectrogram as the input, and we propose a versatile attention module to reduce redundancy and increase efficiency. MHFF-CA is employed to refine intermediate features from WavLM and SSER-VA to improve competitiveness in various tasks, where Contrastive Reconstruction Networks (CRN) and Cross Attention-based Feature Fusion (CAF) are utilized to reconstruct and fuse features, respectively. Effectiveness experiments on IEMOCAP and MELD datasets illustrate MCRVT's superior performance compared to SOTA methods. Exploratory experiments on D-vlog, BioCAS2024, and IEMOCAP (speaker-independent setting) datasets demonstrate the competitiveness of our method in various speech classification tasks.

源语言英语
页(从-至)2189-2199
页数11
期刊IEEE Transactions on Affective Computing
16
3
DOI
出版状态已出版 - 2025
已对外发布

学术指纹

探究 'MCRVT: Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此