摘要
It is a common method to quantify the latent in audio generation and then use diffusion models to estimate noise or data from the corrupted data to generate the quantized latent. Unlike the method, we consider that the targets estimated by the diffusion model include both noise and data, rather than just one of them. Based on this idea and multi-task learning methods, we design the network Dual-Unet, which is simply modified by U-net and can estimate both noise and data simultaneously. Combining Dual-Unet and Variational AutoEncoders with Residual Vector Quantizer, we propose Multitask diffusion model(MTDiffusion), which can generate foley sound audio with a given label. We validate our proposed model on the DCASE task7B dataset. The experimental results show the effectiveness of our proposed model, and both subjective and objective metrics of the generated audio significantly exceed the baseline and the first place ranked on 2023 DCASE task7B.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 486-490 |
| 页数 | 5 |
| 期刊 | Proceedings - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing |
| DOI | |
| 出版状态 | 已出版 - 2024 |
| 已对外发布 | 是 |
| 活动 | 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2024 - Seoul, 韩国 期限: 14 4月 2024 → 19 4月 2024 |
学术指纹
探究 'Mtdiffusion: Multi-Task Diffusion Model With Dual-Unet for Foley Sound Generation' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver