Bi-Spectral Acoustic Features for Robust Speech Recognition

Kazuo ONOE; Shoei SATO; Shinichi HOMMA; Akio KOBAYASHI; Toru IMAI; Tohru TAKAGI

doi:10.1093/ietisy/e91-d.3.631

Special Section on Robust Speech Processing in Realistic Environments

Bi-Spectral Acoustic Features for Robust Speech Recognition

Kazuo ONOE, Shoei SATO, Shinichi HOMMA, Akio KOBAYASHI, Toru IMAI, Tohru TAKAGI

Author information

Keywords: bi-spectrum, non-Gaussianity, phase information, speech recognition

JOURNAL FREE ACCESS

2008 Volume E91.D Issue 3 Pages 631-634

DOI https://doi.org/10.1093/ietisy/e91-d.3.631

Details

Abstract

The extraction of acoustic features for robust speech recognition is very important for improving its performance in realistic environments. The bi-spectrum based on the Fourier transformation of the third-order cumulants expresses the non-Gaussianity and the phase information of the speech signal, showing the dependency between frequency components. In this letter, we propose a method of extracting short-time bispectral acoustic features with averaging features in a single frame. Merged with the conventional Mel frequency cepstral coefficients (MFCC) based on the power spectrum by the principal component analysis (PCA), the proposed features gave a 6.9% relative lower a word error rate in Japanese broadcast news transcription experiments.

Corresponding author

Register with J-STAGE for free!