On using parameterized multi-channel non-causal Wiener filter-adapted convolutional neural networks for distant speech recognition

Citations

SCOPUS

5

초록

Recently, the convolutional neural network (CNN) with multiple microphones was proposed to use the delay-sum (DS) beamformer for distant speech recognition (DSR) and compared to the direct use of multiple acoustic channels as a parallel input to the CNN [1]. We explore the parameterized multi-channel non-causal Wiener filter (PMWF) as the front-end to train the CNN, which is applied to acoustic modeling for DSR. For this, we first present a concise description of the basic PMWF as well as its advantages and then explain how to organize the PMWF into the CNN with a novel architecture. Experimental results on the TIMIT dataset show that the proposed PMWF-based CNN approach outperforms the cross-channel CNN and the DS beamformer when evaluating the word error rate (WER) in various DSR environments.

키워드

beamformingConvolutional neural networksdistant speech recognitionPMWFBandpass filtersBeamformingConvolutionNeural networksAcoustic channelsCausal Wiener filterConvolutional neural networkDistant speech recognitionMultiple microphonesNovel architecturePMWFWord error rateSpeech recognition
제목
On using parameterized multi-channel non-causal Wiener filter-adapted convolutional neural networks for distant speech recognition
저자
Lee, JeehyeChang, Joon-HyukSohn, Jinho
DOI
10.1109/ELINFOCOM.2016.7562963
발행일
2016-09
유형
Conference Paper
저널명
International Conference on Electronics, Information, and Communications, ICEIC 2016
페이지
1 ~ 4