Quad-Net: Melspectrogram Vocoder with Convolutional Layers Restricted by the Quadrature Mirror Filter for Perfect Reconstruction

Citations

SCOPUS

0

초록

Recently, neural vocoders have applied signal processing methods to synthesize speech to reduce computational complexity. However, most methods lack the benefits of a data-driven approach and the flexibility of hyper-parameters, such as filter length, because they rely on fixed signal processing filters. In this paper, we introduce Quad-Net, a network that includes restricted convolutional layers shaped by quadrature mirror synthesis filter banks. It is optimized with a perfect reconstruction loss derived from perfect reconstruction filter banks. This enables us to control filter lengths and degrees of data-drivenness. The results show that the filter parameters trained in our model exhibit characteristics similar to those of other signal processing methods with lower parameters. Furthermore, by increasing the filter length of Quad-Net, we can obtain filters that have complex frequency responses.It shows that a new approach enables the design of more complex filters that are adaptive to neural networks, diverging from previous methods.

키워드

perfect reconstructionquadrature mirror filtersingal processingvocoderAudio signal processingComplex networksComputer visionConvolutionFilter banksFrequency responseMirrorsNeural networksSpeech communication
제목
Quad-Net: Melspectrogram Vocoder with Convolutional Layers Restricted by the Quadrature Mirror Filter for Perfect Reconstruction
저자
Song, Nam-SeokChang, Joon-Hyuk
DOI
10.1109/ICASSP49660.2025.10890659
발행일
2025-03
유형
Conference paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
1 ~ 5