MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS

Miao Liu; Jing Wang; Shicong Li; Fei Xiang; Yue Yao; Lidong Yang

doi:10.1109/ICASSP43922.2022.9747533

MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS

Miao Liu, Jing Wang, Shicong Li, Fei Xiang, Yue Yao, Lidong Yang

信息与电子学院

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

2 引用（Scopus）

摘要

Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognition, but few studies have used it to estimate speech quality. In this paper, we propose a neural-network-based model that splices the deep features extracted by convolutional neural network (CNN) and i-vector on the time axis and uses Transformer encoder as time sequence model. To evaluate the proposed method, we improve the previous prediction models and conduct experiments on Voice Conversion Challenge (VCC) 2018 and 2016 dataset. Results show that i-vector contains information very related to the quality of synthetic speech and the proposed models that utilize i-vector and Transformer encoder highly increase the accuracy of MOSNet and MBNet on both utterance-level and system-level results.

源语言	英语
主期刊名	2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings
出版商	Institute of Electrical and Electronics Engineers Inc.
页	906-910
页数	5
ISBN（电子版）	9781665405409
DOI	https://doi.org/10.1109/ICASSP43922.2022.9747533
出版状态	已出版 - 2022
活动	47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Virtual, Online, 新加坡期限: 23 5月 2022 → 27 5月 2022

出版系列

姓名	ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
卷	2022-May
ISSN（印刷版）	1520-6149

会议

会议	47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022
国家/地区	新加坡
市	Virtual, Online
时期	23/05/22 → 27/05/22

访问文件

10.1109/ICASSP43922.2022.9747533

其它文件与链接

链接到 Scopus 的出版物

引用此

Liu, M., Wang, J., Li, S., Xiang, F., Yao, Y., & Yang, L. (2022). MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS. 在 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings (页码 906-910). (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings; 卷 2022-May). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/ICASSP43922.2022.9747533

Liu, Miao ; Wang, Jing ; Li, Shicong 等. / MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS. 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings. Institute of Electrical and Electronics Engineers Inc., 2022. 页码 906-910 (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings).

@inproceedings{0af07a83c08f45d0b410f7f2f98fcc54,

title = "MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS",

abstract = "Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognition, but few studies have used it to estimate speech quality. In this paper, we propose a neural-network-based model that splices the deep features extracted by convolutional neural network (CNN) and i-vector on the time axis and uses Transformer encoder as time sequence model. To evaluate the proposed method, we improve the previous prediction models and conduct experiments on Voice Conversion Challenge (VCC) 2018 and 2016 dataset. Results show that i-vector contains information very related to the quality of synthetic speech and the proposed models that utilize i-vector and Transformer encoder highly increase the accuracy of MOSNet and MBNet on both utterance-level and system-level results.",

keywords = "Transformer encoder, i-vector, speech quality assessment, speech synthesis",

author = "Miao Liu and Jing Wang and Shicong Li and Fei Xiang and Yue Yao and Lidong Yang",

note = "Publisher Copyright: {\textcopyright} 2022 IEEE; 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 ; Conference date: 23-05-2022 Through 27-05-2022",

year = "2022",

doi = "10.1109/ICASSP43922.2022.9747533",

language = "English",

series = "ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

pages = "906--910",

booktitle = "2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings",

address = "United States",

}

Liu, M, Wang, J, Li, S, Xiang, F, Yao, Y & Yang, L 2022, MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS. 在 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 卷 2022-May, Institute of Electrical and Electronics Engineers Inc., 页码 906-910, 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022, Virtual, Online, 新加坡, 23/05/22. https://doi.org/10.1109/ICASSP43922.2022.9747533

MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS. / Liu, Miao; Wang, Jing; Li, Shicong 等.
2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings. Institute of Electrical and Electronics Engineers Inc., 2022. 页码 906-910 (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings; 卷 2022-May).

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

TY - GEN

T1 - MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS

AU - Liu, Miao

AU - Wang, Jing

AU - Li, Shicong

AU - Xiang, Fei

AU - Yao, Yue

AU - Yang, Lidong

PY - 2022

Y1 - 2022

N2 - Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognition, but few studies have used it to estimate speech quality. In this paper, we propose a neural-network-based model that splices the deep features extracted by convolutional neural network (CNN) and i-vector on the time axis and uses Transformer encoder as time sequence model. To evaluate the proposed method, we improve the previous prediction models and conduct experiments on Voice Conversion Challenge (VCC) 2018 and 2016 dataset. Results show that i-vector contains information very related to the quality of synthetic speech and the proposed models that utilize i-vector and Transformer encoder highly increase the accuracy of MOSNet and MBNet on both utterance-level and system-level results.

AB - Based on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognition, but few studies have used it to estimate speech quality. In this paper, we propose a neural-network-based model that splices the deep features extracted by convolutional neural network (CNN) and i-vector on the time axis and uses Transformer encoder as time sequence model. To evaluate the proposed method, we improve the previous prediction models and conduct experiments on Voice Conversion Challenge (VCC) 2018 and 2016 dataset. Results show that i-vector contains information very related to the quality of synthetic speech and the proposed models that utilize i-vector and Transformer encoder highly increase the accuracy of MOSNet and MBNet on both utterance-level and system-level results.

KW - Transformer encoder

KW - i-vector

KW - speech quality assessment

KW - speech synthesis

UR - http://www.scopus.com/inward/record.url?scp=85131244458&partnerID=8YFLogxK

U2 - 10.1109/ICASSP43922.2022.9747533

DO - 10.1109/ICASSP43922.2022.9747533

M3 - Conference contribution

AN - SCOPUS:85131244458

T3 - ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings

SP - 906

EP - 910

BT - 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings

PB - Institute of Electrical and Electronics Engineers Inc.

T2 - 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022

Y2 - 23 May 2022 through 27 May 2022

ER -

Liu M, Wang J, Li S, Xiang F, Yao Y, Yang L. MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS. 在 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - Proceedings. Institute of Electrical and Electronics Engineers Inc. 2022. 页码 906-910. (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings). doi: 10.1109/ICASSP43922.2022.9747533

MOS PREDICTOR FOR SYNTHETIC SPEECH WITH I-VECTOR INPUTS

摘要

出版系列

会议

访问文件

其它文件与链接

指纹

引用此