Skip to main navigation Skip to search Skip to main content

Dual-Stage Cross-Modal Attention Network for Multimodal Fake News Detection

  • Beijing Institute of Technology
  • University of Chinese Academy of Sciences

Research output: Contribution to journalConference articlepeer-review

Abstract

Detecting fake news has become increasingly challenging in the era of multimodal social media, where deceptive content often combines misleading text with incongruent images. Existing methods frequently suffer from three key limitations: (1) misaligned semantic representations across modalities, (2) reliance on overly complex or inefficient fusion mechanisms, and (3) often overlook cross-modal semantic (in)consistencies, which are widely regarded as critical cues for detecting fake news. To address these challenges, we propose the Cross-Modal Aligned Attention Network (CMAAN), an efficient framework tailored for multimodal rumor detection. CMAAN utilizes CLIP to obtain semantically aligned text and image embeddings, which are then integrated into a compact joint representation via a lightweight gating-and-weighting module. Central to our approach is a dual-stage cross-modal attention mechanism: the first stage refines unimodal features under the guidance of the fused representation, while the second facilitates bidirectional interactions to enable effective cross-modal reasoning. Experiments on two widely used real-world multimodal rumor detection datasets demonstrate the effectiveness of the proposed approach.

Original languageEnglish
Pages (from-to)593-598
Number of pages6
JournalProceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD
Issue number2026
DOIs
Publication statusPublished - 2026
Event29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China
Duration: 13 May 202615 May 2026

Keywords

  • CLIP
  • fake news detection
  • multimodal
  • social media

Fingerprint

Dive into the research topics of 'Dual-Stage Cross-Modal Attention Network for Multimodal Fake News Detection'. Together they form a unique fingerprint.

Cite this