Skip to main navigation Skip to search Skip to main content

Cross-Attention Audio–Visual Fusion Based on Multi-Scale Vision Transformers for Emotion Recognition

  • Chengao Bao
  • , Luefeng Chen*
  • , Min Li
  • , Min Wu
  • , Witold Pedrycz
  • , Kaoru Hirota
  • *Corresponding author for this work
  • China University of Geosciences, Wuhan
  • University of Alberta
  • Institute of Science Tokyo

Research output: Contribution to journalArticlepeer-review

Abstract

A multimodal emotion recognition method based on a multi-scale vision Transformer and cross attention mechanism (MSFCA) is proposed. The proposed method integrates a multi-scale vision Transformer and a cross-attention mechanism to fully exploit the complementary information between facial expressions and speech modalities, thereby enabling more effective feature extraction and fusion. By incorporating L1 regularization, the sparsity and robustness of the fused features are enhanced, thereby improving the accuracy and generalization capability of multi-modal emotion recognition. Experimental results on the eNTERFACE’05 and RAVDESS multi-modal emotion recognition datasets demonstrate that the proposed MSFCA method achieves recognition accuracies of 86.16% and 87.14%, respectively, which satisfy the requirements for reliable multi-modal emotion recognition in practical applications.

Original languageEnglish
Pages (from-to)749-760
Number of pages12
JournalJournal of Advanced Computational Intelligence and Intelligent Informatics
Volume30
Issue number3
DOIs
Publication statusPublished - May 2026
Externally publishedYes

Keywords

  • cross-attention
  • multi-scale features
  • multimodal emotion recognition

Fingerprint

Dive into the research topics of 'Cross-Attention Audio–Visual Fusion Based on Multi-Scale Vision Transformers for Emotion Recognition'. Together they form a unique fingerprint.

Cite this