TY - GEN
T1 - Partially Fake Audio Detection Based on Mamba and Tensor Feature Fusion
AU - Liu, Hanyue
AU - Liu, Miao
AU - Deng, Mengyuan
AU - Wei, Fangda
AU - Lang, Yue
AU - Zhao, Shenghui
AU - Wang, Jing
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In recent years, models related to Artificial Intelligence Generated Content (AIGC) have advanced rapidly, opening up new possibilities for generating realistic speech. However, if misused, this technology can pose significant risks to information security. Consequently, the task of deepfake audio detection has emerged. Within this domain, the Manipulation Region Location (MRL) task specifically aims to identify the manipulated segments of speech, offering greater precision and facilitating downstream tasks such as intent analysis. In this paper, we first investigate the performance of various feature types in the MRL task to determine which features exhibit the best generalization ability. Then we propose a tensor-based feature fusion strategy to effectively capture the interrelationships among different features and produce a more representative fused feature. Furthermore, leveraging the strong temporal modeling capabilities of Mamba, we incorporate it into our framework. To the best of our knowledge, this is the first work to introduce Mamba into the MRL task. The model is trained on the ADD2023 dataset. Experimental results on the test set demonstrate that our approach achieves a 37.5% improvement in performance over the baseline system and outperforms other compared models.
AB - In recent years, models related to Artificial Intelligence Generated Content (AIGC) have advanced rapidly, opening up new possibilities for generating realistic speech. However, if misused, this technology can pose significant risks to information security. Consequently, the task of deepfake audio detection has emerged. Within this domain, the Manipulation Region Location (MRL) task specifically aims to identify the manipulated segments of speech, offering greater precision and facilitating downstream tasks such as intent analysis. In this paper, we first investigate the performance of various feature types in the MRL task to determine which features exhibit the best generalization ability. Then we propose a tensor-based feature fusion strategy to effectively capture the interrelationships among different features and produce a more representative fused feature. Furthermore, leveraging the strong temporal modeling capabilities of Mamba, we incorporate it into our framework. To the best of our knowledge, this is the first work to introduce Mamba into the MRL task. The model is trained on the ADD2023 dataset. Experimental results on the test set demonstrate that our approach achieves a 37.5% improvement in performance over the baseline system and outperforms other compared models.
KW - Mamba
KW - feature fusion
KW - manipulation region location
KW - partially fake audio detection
KW - tensor network
UR - https://www.scopus.com/pages/publications/105043444005
U2 - 10.1109/ACAIT67930.2025.11521894
DO - 10.1109/ACAIT67930.2025.11521894
M3 - Conference contribution
AN - SCOPUS:105043444005
T3 - Proceedings of 2025 9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025
SP - 1059
EP - 1063
BT - Proceedings of 2025 9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025
Y2 - 12 September 2025 through 14 September 2025
ER -