Skip to main navigation Skip to search Skip to main content

Enhanced Prompt-Supervised Multimodal Depression Detection Based on Social Media

  • Yongfeng Tao
  • , Zilin Guo
  • , Zhichao Yang
  • , Bin Hu
  • , Minqiang Yang*
  • *Corresponding author for this work
  • Lanzhou University
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Depression detection refers to the task of automatically identifying or estimating depressive states and their severity from multimodal data, including behavioral, linguistic, acoustic, and visual cues. Among these modalities, nonverbal behavioral data, such as visual and acoustic signals, has received particular attention because it provides objective indicators of depressive tendencies. However, neglecting critical auxiliary enhancement descriptors poses challenges in distinguishing between depressed and non-depressed states, especially in low-quality data. To tackle this challenge, we introduce EPMdd, an Enhanced Prompt-supervised Multimodal depression detection framework. This approach leverages a Probability-Based Enhancement Module (PBEM) that employs a pre-trained text-based Large Language Model (LLM) as a behavioral analysis engine to integrate multimodal nonverbal behavioral cues with linguistic descriptions during training. PBEM enables robust and fine-grained detection of depressive emotional variations, strengthening the model’s feature extraction capability while maintaining computational efficiency. To further capture temporal dependencies and enhance discriminative representation, we design a Local-Global Self-Attention Module (LGAM) that jointly learns local and global contextual features. More importantly, we propose a Multimodal Cross-Attention Fusion (MCAF) module that fuses high-level semantic representations derived from low-level visual and acoustic features, facilitating comprehensive spatiotemporal feature learning. Extensive experiments on two large-scale public social-media datasets, D-Vlog and LMVD, demonstrate that EPMdd achieves competitive performance, particularly excelling in precision with 74.07% and 77.77% respectively, showcasing its significant advantages in relevant detection tasks.

Original languageEnglish
JournalIEEE Transactions on Consumer Electronics
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Depression Detection
  • Large Language Model
  • Multimodal Fusion
  • Social Media Vlogs

Fingerprint

Dive into the research topics of 'Enhanced Prompt-Supervised Multimodal Depression Detection Based on Social Media'. Together they form a unique fingerprint.

Cite this