跳到主要导航 跳到搜索 跳到主要内容

Multimodal Emotion Recognition Based on Multi-Scale Facial Features and Cross-Modal Attention

  • Chengao Bao
  • , Luefeng Chen*
  • , Min Li
  • , Min Wu
  • , Witold Pedrycz
  • , Kaoru Hirota
  • *此作品的通讯作者
  • China University of Geosciences, Wuhan
  • University of Alberta
  • Institute of Science Tokyo

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

A multi-modal emotion recognition method based on facial multi-scale features and cross-modal attention (MS-FCA) network is proposed. The MSFCA model improves the traditional single-branch ViT network into a two-branch ViT architecture by using classification tokens in each branch to interact with picture embeddings in the other branch, which facilitates effective interactions between different scales of information. Subsequently, audio features are extracted using ResNet18 network. The cross-modal attention mechanism is used to obtain the weight matrices between different modal features, making full use of inter-modal correlation and effectively fusing visual and audio features for more accurate emotion recognition. Two datasets are used for the experiments: eNTERFACE'05 and REDVESS dataset. The experimental results show that the accuracy of the proposed method on the eNTERFACE'05 and REDVESS datasets is 85.42% and 83.84% respectively, which proves the effectiveness of the proposed method.

源语言英语
主期刊名2025 International Conference on Industrial Technology, ICIT 2025 - Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798331521950
DOI
出版状态已出版 - 2025
已对外发布
活动26th International Conference on Industrial Technology, ICIT 2025 - Wuhan, 中国
期限: 26 3月 202528 3月 2025

出版系列

姓名Proceedings of the IEEE International Conference on Industrial Technology
ISSN(印刷版)2641-0184
ISSN(电子版)2643-2978

会议

会议26th International Conference on Industrial Technology, ICIT 2025
国家/地区中国
Wuhan
时期26/03/2528/03/25

指纹

探究 'Multimodal Emotion Recognition Based on Multi-Scale Facial Features and Cross-Modal Attention' 的科研主题。它们共同构成独一无二的指纹。

引用此