跳到主要导航 跳到搜索 跳到主要内容

Mtdiffusion: Multi-Task Diffusion Model With Dual-Unet for Foley Sound Generation

  • Beijing Institute of Technology

科研成果: 期刊稿件会议文章同行评审

摘要

It is a common method to quantify the latent in audio generation and then use diffusion models to estimate noise or data from the corrupted data to generate the quantized latent. Unlike the method, we consider that the targets estimated by the diffusion model include both noise and data, rather than just one of them. Based on this idea and multi-task learning methods, we design the network Dual-Unet, which is simply modified by U-net and can estimate both noise and data simultaneously. Combining Dual-Unet and Variational AutoEncoders with Residual Vector Quantizer, we propose Multitask diffusion model(MTDiffusion), which can generate foley sound audio with a given label. We validate our proposed model on the DCASE task7B dataset. The experimental results show the effectiveness of our proposed model, and both subjective and objective metrics of the generated audio significantly exceed the baseline and the first place ranked on 2023 DCASE task7B.

源语言英语
页(从-至)486-490
页数5
期刊Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing
DOI
出版状态已出版 - 2024
已对外发布
活动2024 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2024 - Seoul, 韩国
期限: 14 4月 202419 4月 2024

学术指纹

探究 'Mtdiffusion: Multi-Task Diffusion Model With Dual-Unet for Foley Sound Generation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此