상세 보기
Neural ATSM: Fully Neural Network-based Adaptive Time-Scale Modification Using Sentence-Specific Dynamic Control
- Lee, Jaeuk;
- Jang, Sohee;
- Chang, Joon-Hyuk
WEB OF SCIENCE
0SCOPUS
0초록
Adaptive time-scale modification (ATSM) adaptively adjusts audio speed and improves upon previous systems by tailoring the scale for each phoneme in two steps: phoneme positioning via Montreal forced aligner (MFA) and reconstruction with adaptive speaking rate. However, ATSM's phoneme-specific rate is constant regardless of sentences, and MFA struggles with precise phoneme alignment in synthetic speech. Driven by this, we propose a fully neural networks-based ATSM (Neural ATSM) that dynamically controls each phoneme's speaking rate to vary from sentence to sentence. It predicts phoneme-level rates using a speaking rate predictor and flexibly modifies the scales to fit sentence context using Gaussian upsampling and attention mechanism, ensuring feature similarity with Soft-dynamic time warping (DTW) loss. We also integrate a variational autoencoder (VAE) and flow models for enhanced time-scaled signals. Experimental results show that Neural ATSM outperforms ATSM for real and synthesized speech.
키워드
- 제목
- Neural ATSM: Fully Neural Network-based Adaptive Time-Scale Modification Using Sentence-Specific Dynamic Control
- 저자
- Lee, Jaeuk; Jang, Sohee; Chang, Joon-Hyuk
- 발행일
- 2024-09
- 유형
- Proceedings Paper
- 저널명
- INTERSPEECH 2024
- 페이지
- 4903 ~ 4907