Abstract
Two challenges for Speech Emotion Recognition (SER) are efficiently capturing features to explore speech-emotion correlations and reducing the SER model redundancy while increasing performance. To address these challenges, we propose Multi-Hierarchical Cross-Reconstruction Networks with Versatile Transformer (MCRVT), including a Spectrogram-Based SER Model with Versatile Attention (SSER-VA), a pre-trained model (WavLM), and Multi-Hierarchical Feature Fusion with Cross Attention (MHFF-CA). Specifically, SSER-VA takes the log-Mel spectrogram as the input, and we propose a versatile attention module to reduce redundancy and increase efficiency. MHFF-CA is employed to refine intermediate features from WavLM and SSER-VA to improve competitiveness in various tasks, where Contrastive Reconstruction Networks (CRN) and Cross Attention-based Feature Fusion (CAF) are utilized to reconstruct and fuse features, respectively. Effectiveness experiments on IEMOCAP and MELD datasets illustrate MCRVT's superior performance compared to SOTA methods. Exploratory experiments on D-vlog, BioCAS2024, and IEMOCAP (speaker-independent setting) datasets demonstrate the competitiveness of our method in various speech classification tasks.
| Original language | English |
|---|---|
| Pages (from-to) | 2189-2199 |
| Number of pages | 11 |
| Journal | IEEE Transactions on Affective Computing |
| Volume | 16 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - 2025 |
| Externally published | Yes |
Keywords
- Speech emotion recognition
- cross-reconstruction networks
- speech depression recognition
- transformer
- versatile attention
Fingerprint
Dive into the research topics of 'MCRVT: Multi-Hierarchical Cross-Reconstruction Networks With Versatile Transformer for Speech Emotion Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver