Abstract
A multimodal emotion recognition method based on a multi-scale vision Transformer and cross attention mechanism (MSFCA) is proposed. The proposed method integrates a multi-scale vision Transformer and a cross-attention mechanism to fully exploit the complementary information between facial expressions and speech modalities, thereby enabling more effective feature extraction and fusion. By incorporating L1 regularization, the sparsity and robustness of the fused features are enhanced, thereby improving the accuracy and generalization capability of multi-modal emotion recognition. Experimental results on the eNTERFACE’05 and RAVDESS multi-modal emotion recognition datasets demonstrate that the proposed MSFCA method achieves recognition accuracies of 86.16% and 87.14%, respectively, which satisfy the requirements for reliable multi-modal emotion recognition in practical applications.
| Original language | English |
|---|---|
| Pages (from-to) | 749-760 |
| Number of pages | 12 |
| Journal | Journal of Advanced Computational Intelligence and Intelligent Informatics |
| Volume | 30 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - May 2026 |
| Externally published | Yes |
Keywords
- cross-attention
- multi-scale features
- multimodal emotion recognition
Fingerprint
Dive into the research topics of 'Cross-Attention Audio–Visual Fusion Based on Multi-Scale Vision Transformers for Emotion Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver