CNN을 위한 사투리 음성 데이터의 Spectrogram 이미지 변환 적용의 POC 검증

POC Analysis of Spectrogram Image Transformation Application to Dialect Speech Data for CNN

초록

In essence, audio data is time-series data. Therefore, for audio classification, time-series algorithms such as ARIMA (Autoregressive integrated moving average), ES (Exponential smoothing), or RNN (Recurrent neural network) in machine learning terms are commonly employed. Another method is to use a spectrogram as an image representing audio data rather than a time series number array as the input to the CNN(Convolutional neural network) learning process instead of RNN. In this paper, we use spectrogram analysis and mel-spectrogram analysis images of audio data as input data, using the CNN technique instead of RNN for time series data analysis, and we propose a CNN model to extract patterns of dialect audio. In addition, the feasibility was verified by performing a POC(Proof of concept) using a Python-based prototype of the proposed model.

키워드

Dialect audio dataSpectrogramMel-spectrogramCNNRNN
제목
CNN을 위한 사투리 음성 데이터의 Spectrogram 이미지 변환 적용의 POC 검증
제목 (타언어)
POC Analysis of Spectrogram Image Transformation Application to Dialect Speech Data for CNN
저자
차병래권용
DOI
10.22733/JITAE.2024.14.02.001
발행일
2024-08
저널명
Journal of Information Technology and Applied Engineering
14
2
페이지
1 ~ 7