Skip to main navigation Skip to search Skip to main content

The formant structure based feature parameter for speech recognition

  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we proposed a new set of speech feature parameters based on formant structure information. The speech signal is first divided into a gammatone filterbank, and then the Teager energy signal of each sub-band is extracted according to Teager Energy Operator. Two different energy separation algorithms, DESA-1 and DESA-2, are applied for obtain the instantaneous amplitude and frequency envelope of the formants, respectively. Finally, the feature vector is constructed by the amplitude and frequency information of the formants. The motivation of developing this new feature is that the formant location information is a quite distinct speech representation but seldom applied into speech recognition system before, and the conventional feature parameters such as mel-frequency cepstral coefficients (MFCC) and linear prediction cepstral coefficients (LPCC) do not explicitly model spectral peak information which is very important clue to identify the different phones. A Mandarin digit string recognition task is performed for evaluating the performance of the proposed feature parameter. The recognition results show an improved speech recognition performance compared to the conventional MFCC and LPCC.

Original languageEnglish
Title of host publicationProceedings of the 2003 IEEE Workshop on Statistical Signal Processing, SSP 2003
PublisherIEEE Computer Society
Pages605-608
Number of pages4
ISBN (Electronic)0780379977
DOIs
Publication statusPublished - 2003
EventIEEE Workshop on Statistical Signal Processing, SSP 2003 - St. Louis, United States
Duration: 28 Sept 20031 Oct 2003

Publication series

NameIEEE Workshop on Statistical Signal Processing Proceedings
Volume2003-January

Conference

ConferenceIEEE Workshop on Statistical Signal Processing, SSP 2003
Country/TerritoryUnited States
CitySt. Louis
Period28/09/031/10/03

Keywords

  • Automatic speech recognition
  • Cepstral analysis
  • Data mining
  • Frequency estimation
  • Frequency modulation
  • Mel frequency cepstral coefficient
  • Resonance
  • Signal processing
  • Speech processing
  • Speech recognition

Fingerprint

Dive into the research topics of 'The formant structure based feature parameter for speech recognition'. Together they form a unique fingerprint.

Cite this