상세 보기
FiLM Conditioning with Enhanced Feature to the Transformer-based End-to-End Noisy Speech Recognition
- Yang, Da-Hee;
- Chang, Joon-Hyuk
WEB OF SCIENCE
2SCOPUS
1초록
Ensuring robustness against environmental noise is an important concern in the design of automatic speech recognition (ASR) systems. This is typically achieved by utilizing a speech enhancement (SE) network in an ASR system to boost noise robustness. The performance of ASR systems can be improved using SE networks as a front-end or by retraining the ASR system on enhanced speech. Although the SE network is effective, it does not always result in improved performance in the ASR system owing to artifacts. To address this problem, we propose the use of enhanced speech from an SE network as a conditioning feature instead of a direct input feature of the ASR system. This is achieved by stacking a feature-wise linear modulation (FiLM) layer on each transformer layer of the end-to-end ASR encoder and combining the input and conditioning features. The results indicate that the proposed FiLM training method exhibits greater robustness against noise owing to the use of enhanced speech as conditioning information rather than as direct ASR input.
키워드
- 제목
- FiLM Conditioning with Enhanced Feature to the Transformer-based End-to-End Noisy Speech Recognition
- 저자
- Yang, Da-Hee; Chang, Joon-Hyuk
- 발행일
- 2022-09
- 유형
- Proceedings Paper
- 저널명
- INTERSPEECH 2022
- 권
- 2022-September
- 페이지
- 4098 ~ 4102