Skip to main navigation Skip to search Skip to main content

Partially Fake Audio Detection Based on Mamba and Tensor Feature Fusion

  • Hanyue Liu
  • , Miao Liu
  • , Mengyuan Deng
  • , Fangda Wei
  • , Yue Lang
  • , Shenghui Zhao
  • , Jing Wang*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Hebei North University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In recent years, models related to Artificial Intelligence Generated Content (AIGC) have advanced rapidly, opening up new possibilities for generating realistic speech. However, if misused, this technology can pose significant risks to information security. Consequently, the task of deepfake audio detection has emerged. Within this domain, the Manipulation Region Location (MRL) task specifically aims to identify the manipulated segments of speech, offering greater precision and facilitating downstream tasks such as intent analysis. In this paper, we first investigate the performance of various feature types in the MRL task to determine which features exhibit the best generalization ability. Then we propose a tensor-based feature fusion strategy to effectively capture the interrelationships among different features and produce a more representative fused feature. Furthermore, leveraging the strong temporal modeling capabilities of Mamba, we incorporate it into our framework. To the best of our knowledge, this is the first work to introduce Mamba into the MRL task. The model is trained on the ADD2023 dataset. Experimental results on the test set demonstrate that our approach achieves a 37.5% improvement in performance over the baseline system and outperforms other compared models.

Original languageEnglish
Title of host publicationProceedings of 2025 9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1059-1063
Number of pages5
ISBN (Electronic)9798331587871
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025 - Ordos, China
Duration: 12 Sept 202514 Sept 2025

Publication series

NameProceedings of 2025 9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025

Conference

Conference9th Asian Conference on Artificial Intelligence Technology, ACAIT 2025
Country/TerritoryChina
CityOrdos
Period12/09/2514/09/25

Keywords

  • Mamba
  • feature fusion
  • manipulation region location
  • partially fake audio detection
  • tensor network

Fingerprint

Dive into the research topics of 'Partially Fake Audio Detection Based on Mamba and Tensor Feature Fusion'. Together they form a unique fingerprint.

Cite this