跳到主要导航 跳到搜索 跳到主要内容

Neural Audio Coding with Deep Complex Networks

  • Jiawei Ru
  • , Lizhong Wang
  • , Maoshen Jia
  • , Liang Wen
  • , Chunxi Wang
  • , Yuhao Zhao
  • , Jing Wang
  • Beijing University of Technology
  • Samsung R&D Institute China

科研成果: 期刊稿件会议文章同行评审

摘要

This paper proposes a transform domain audio coding method based on deep complex networks. In the proposed codec, the time-frequency spectrum of the audio signal is fed to the encoder which consists of complex convolutional blocks and a frequency-temporal modeling module to obtain the extracted features which are then quantized with a target bitrate by the vector quantizer. The structure of the decoder which reconstruct the time-frequency spectrum of the audio from quantized features is symmetrical to the encoder. In this paper, a structure combining the complex multi-head self-attention module and the complex long short-term memory is proposed to capture both frequency and temporal dependencies. Subjective and objective evaluation tests show the advantage of the proposed method.

源语言英语
文章编号012005
期刊Journal of Physics: Conference Series
2759
1
DOI
出版状态已出版 - 2024
活动2024 8th International Conference on Machine Vision and Information Technology, CMVIT 2024 - Hybrid, Singapore, 新加坡
期限: 23 2月 202425 2月 2024

指纹

探究 'Neural Audio Coding with Deep Complex Networks' 的科研主题。它们共同构成独一无二的指纹。

引用此