跳到主要导航 跳到搜索 跳到主要内容

Cross-Attention Audio–Visual Fusion Based on Multi-Scale Vision Transformers for Emotion Recognition

  • Chengao Bao
  • , Luefeng Chen*
  • , Min Li
  • , Min Wu
  • , Witold Pedrycz
  • , Kaoru Hirota
  • *此作品的通讯作者
  • China University of Geosciences, Wuhan
  • University of Alberta
  • Institute of Science Tokyo

科研成果: 期刊稿件文章同行评审

摘要

A multimodal emotion recognition method based on a multi-scale vision Transformer and cross attention mechanism (MSFCA) is proposed. The proposed method integrates a multi-scale vision Transformer and a cross-attention mechanism to fully exploit the complementary information between facial expressions and speech modalities, thereby enabling more effective feature extraction and fusion. By incorporating L1 regularization, the sparsity and robustness of the fused features are enhanced, thereby improving the accuracy and generalization capability of multi-modal emotion recognition. Experimental results on the eNTERFACE’05 and RAVDESS multi-modal emotion recognition datasets demonstrate that the proposed MSFCA method achieves recognition accuracies of 86.16% and 87.14%, respectively, which satisfy the requirements for reliable multi-modal emotion recognition in practical applications.

源语言英语
页(从-至)749-760
页数12
期刊Journal of Advanced Computational Intelligence and Intelligent Informatics
30
3
DOI
出版状态已出版 - 5月 2026
已对外发布

指纹

探究 'Cross-Attention Audio–Visual Fusion Based on Multi-Scale Vision Transformers for Emotion Recognition' 的科研主题。它们共同构成独一无二的指纹。

引用此