TY - JOUR
T1 - MM-LLM for Depression
T2 - Towards an Augmented Intelligent Diagnosis and Intervention by Fusing Multimodal Data and Large Language Models
AU - Shen, Jian
AU - Ma, Yu
AU - Gao, Haoran
AU - Lu, Chenyang
AU - Ma, Ruirui
AU - Zhou, Xinnan
AU - Xu, Wentian
AU - Wang, Jiayue
AU - Li, Changlong
AU - Ding, Yiwen
AU - Wang, Ran
AU - An, Cuixia
AU - Zhang, Yanan
AU - Xu, Chen
AU - Hu, Bin
N1 - Publisher Copyright:
© 2010-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Depression is a highly prevalent mental disorder that remains difficult to assess and manage using traditional questionnaire- or interview-based approaches due to subjectivity and limited real-time adaptability. Existing support approaches further lack personalized interaction and adaptive feedback, limiting timely and individualized mental health support. To address these challenges, we propose MM-LLM, an augmented intelligent screening, assessment-support, and intervention-oriented response-generation framework by fusing multimodal data and Large Language Models (LLM). First, a cross-modal guidance module integrates EEG with speech and text representations using pretrained models to enhance neural discriminability. Second, a cross-domain knowledge transfer strategy aligns semantic spaces across subjects and task paradigms, enabling personalized yet generalizable modeling of depression-related features. Third, an LLM-based response-support module leverages state tracking, knowledge-graph retrieval, and retrieval-augmented generation to generate individualized supportive responses within a turn-level feedback cycle. Experimental results provide a proof-of-concept for MM-LLM's ability to enhance recognition accuracy and support response generation, while the small pilot pre-post evaluation provides only a preliminary short-term symptom-change signal that requires validation in larger controlled studies.
AB - Depression is a highly prevalent mental disorder that remains difficult to assess and manage using traditional questionnaire- or interview-based approaches due to subjectivity and limited real-time adaptability. Existing support approaches further lack personalized interaction and adaptive feedback, limiting timely and individualized mental health support. To address these challenges, we propose MM-LLM, an augmented intelligent screening, assessment-support, and intervention-oriented response-generation framework by fusing multimodal data and Large Language Models (LLM). First, a cross-modal guidance module integrates EEG with speech and text representations using pretrained models to enhance neural discriminability. Second, a cross-domain knowledge transfer strategy aligns semantic spaces across subjects and task paradigms, enabling personalized yet generalizable modeling of depression-related features. Third, an LLM-based response-support module leverages state tracking, knowledge-graph retrieval, and retrieval-augmented generation to generate individualized supportive responses within a turn-level feedback cycle. Experimental results provide a proof-of-concept for MM-LLM's ability to enhance recognition accuracy and support response generation, while the small pilot pre-post evaluation provides only a preliminary short-term symptom-change signal that requires validation in larger controlled studies.
KW - Cross-domain knowledge transfer
KW - Depression recognition
KW - Large language models (LLM)
KW - Multimodal fusion
KW - Personalized response support
UR - https://www.scopus.com/pages/publications/105043698183
U2 - 10.1109/TAFFC.2026.3708457
DO - 10.1109/TAFFC.2026.3708457
M3 - Article
AN - SCOPUS:105043698183
SN - 1949-3045
JO - IEEE Transactions on Affective Computing
JF - IEEE Transactions on Affective Computing
ER -