Abstract
Detecting fake news has become increasingly challenging in the era of multimodal social media, where deceptive content often combines misleading text with incongruent images. Existing methods frequently suffer from three key limitations: (1) misaligned semantic representations across modalities, (2) reliance on overly complex or inefficient fusion mechanisms, and (3) often overlook cross-modal semantic (in)consistencies, which are widely regarded as critical cues for detecting fake news. To address these challenges, we propose the Cross-Modal Aligned Attention Network (CMAAN), an efficient framework tailored for multimodal rumor detection. CMAAN utilizes CLIP to obtain semantically aligned text and image embeddings, which are then integrated into a compact joint representation via a lightweight gating-and-weighting module. Central to our approach is a dual-stage cross-modal attention mechanism: the first stage refines unimodal features under the guidance of the fused representation, while the second facilitates bidirectional interactions to enable effective cross-modal reasoning. Experiments on two widely used real-world multimodal rumor detection datasets demonstrate the effectiveness of the proposed approach.
| Original language | English |
|---|---|
| Pages (from-to) | 593-598 |
| Number of pages | 6 |
| Journal | Proceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD |
| Issue number | 2026 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China Duration: 13 May 2026 → 15 May 2026 |
Keywords
- CLIP
- fake news detection
- multimodal
- social media
Fingerprint
Dive into the research topics of 'Dual-Stage Cross-Modal Attention Network for Multimodal Fake News Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver