TY - GEN
T1 - The formant structure based feature parameter for speech recognition
AU - Zhao, Junhui
AU - Kuang, Jingming
AU - Xie, Xiang
N1 - Publisher Copyright:
© 2003 IEEE.
PY - 2003
Y1 - 2003
N2 - In this paper, we proposed a new set of speech feature parameters based on formant structure information. The speech signal is first divided into a gammatone filterbank, and then the Teager energy signal of each sub-band is extracted according to Teager Energy Operator. Two different energy separation algorithms, DESA-1 and DESA-2, are applied for obtain the instantaneous amplitude and frequency envelope of the formants, respectively. Finally, the feature vector is constructed by the amplitude and frequency information of the formants. The motivation of developing this new feature is that the formant location information is a quite distinct speech representation but seldom applied into speech recognition system before, and the conventional feature parameters such as mel-frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC) do not explicitly model spectral peak information which is very important clue to identify the different phones. A Mandarin digit string recognition task is performed for evaluating the performance of the proposed feature parameter. The recognition results show an improved speech recognition performance compared to the conventional MFCC and LPCC.
AB - In this paper, we proposed a new set of speech feature parameters based on formant structure information. The speech signal is first divided into a gammatone filterbank, and then the Teager energy signal of each sub-band is extracted according to Teager Energy Operator. Two different energy separation algorithms, DESA-1 and DESA-2, are applied for obtain the instantaneous amplitude and frequency envelope of the formants, respectively. Finally, the feature vector is constructed by the amplitude and frequency information of the formants. The motivation of developing this new feature is that the formant location information is a quite distinct speech representation but seldom applied into speech recognition system before, and the conventional feature parameters such as mel-frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC) do not explicitly model spectral peak information which is very important clue to identify the different phones. A Mandarin digit string recognition task is performed for evaluating the performance of the proposed feature parameter. The recognition results show an improved speech recognition performance compared to the conventional MFCC and LPCC.
KW - Automatic speech recognition
KW - Cepstral analysis
KW - Data mining
KW - Frequency estimation
KW - Frequency modulation
KW - Mel frequency cepstral coefficient
KW - Resonance
KW - Signal processing
KW - Speech processing
KW - Speech recognition
UR - https://www.scopus.com/pages/publications/84947601505
U2 - 10.1109/SSP.2003.1289551
DO - 10.1109/SSP.2003.1289551
M3 - Conference contribution
AN - SCOPUS:84947601505
T3 - IEEE Workshop on Statistical Signal Processing Proceedings
SP - 605
EP - 608
BT - Proceedings of the 2003 IEEE Workshop on Statistical Signal Processing, SSP 2003
PB - IEEE Computer Society
T2 - IEEE Workshop on Statistical Signal Processing, SSP 2003
Y2 - 28 September 2003 through 1 October 2003
ER -