TY - GEN
T1 - A Method for Anomaly Detection in Surveillance Video Based on Multimodal Large Language Model
AU - Zhou, Chongqin
AU - Mersha, Bemnet Wondimagegnehu
AU - Dai, Yaping
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - With the continuous improvement of urban security systems, intelligent monitoring technology is being widely applied in various fields of social governance. Anomaly detection in surveillance videos plays a crucial role in the development of smart cities. With breakthroughs in crossmodal understanding and reasoning using multimodal large language models, their application in anomaly detection has become a promising new approach. However, directly applying multimodal large language models to surveillance scenarios for anomaly detection using prompt engineering still faces challenges such as significant domain bias in surveillance videos and insufficient output interpretability. To address these challenges, we propose a method called Lo-CoT based on adaptive LoRa fine-tuning and CoT supervised fine-tuning within the surveillance domain, significantly improving the understanding of fine-grained behaviors and the interpretability of anomalies in surveillance scenarios and the accuracy of anomaly detection. We applied our Lo-CoT method to the MSAD dataset and compared it with others' studies, achieving an accuracy improvement of nearly 9 percentage points.
AB - With the continuous improvement of urban security systems, intelligent monitoring technology is being widely applied in various fields of social governance. Anomaly detection in surveillance videos plays a crucial role in the development of smart cities. With breakthroughs in crossmodal understanding and reasoning using multimodal large language models, their application in anomaly detection has become a promising new approach. However, directly applying multimodal large language models to surveillance scenarios for anomaly detection using prompt engineering still faces challenges such as significant domain bias in surveillance videos and insufficient output interpretability. To address these challenges, we propose a method called Lo-CoT based on adaptive LoRa fine-tuning and CoT supervised fine-tuning within the surveillance domain, significantly improving the understanding of fine-grained behaviors and the interpretability of anomalies in surveillance scenarios and the accuracy of anomaly detection. We applied our Lo-CoT method to the MSAD dataset and compared it with others' studies, achieving an accuracy improvement of nearly 9 percentage points.
KW - anomaly detection
KW - multimodal large language models
KW - smart cities
KW - surveillance video
UR - https://www.scopus.com/pages/publications/105043918014
U2 - 10.1109/CCDC69976.2026.11560845
DO - 10.1109/CCDC69976.2026.11560845
M3 - Conference contribution
AN - SCOPUS:105043918014
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 1970
EP - 1975
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -