Abstract
Depression is a prevalent psychiatric condition, affecting nearly 280 million individuals worldwide. Traditional approaches to diagnosing and treating depression are often highly subjective, lacking objective and effective methods for accurate recognition and non-pharmacological interventions. To address these challenges, we propose a novel framework for depression recognition and intervention that integrates multimodal physiological signals with large language models (LLMs). First, we employ LLMs to perform high-fidelity reconstruction and high-order feature generation on multimodal physiological signals, thereby effectively recovering missing or noisy information from the raw data. Next, we design a multimodal alignment module that leverages Cross-Modal Attention (CMA) and LLMs architectures to enable deep fusion of heterogeneous physiological data. During the intervention phase, additional modalities such as text and speech are incorporated to enhance user interaction, enabling personalized and intelligent monitoring of dynamic psychological states and the generation of closed-loop, negative-feedback-based intervention strategies. Through comprehensive multimodal data collection and systematic experimental validation, our results demonstrate that the proposed method significantly outperforms traditional approaches in terms of signal reconstruction quality, depression recognition accuracy, and intervention efficacy. This study offers a novel technical pathway and theoretical foundation for the early detection and intelligent intervention of depression, demonstrating substantial scientific value and application potential.
| Original language | English |
|---|---|
| Article number | 103772 |
| Journal | Information Fusion |
| Volume | 127 |
| DOIs | |
| Publication status | Published - Mar 2026 |
Keywords
- Alignment
- Depression recognition and intervention
- LLMs
- Multimodal physiological signals
Fingerprint
Dive into the research topics of 'EmoSavior: Depression recognition and intervention via multimodal physiological signals and large language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver