Diffusion-based Target Device Style Transfer for Robust Acoustic Scene Classification

Citations

SCOPUS

0

초록

Audio signal processing systems often operate differently depending on the recording devices, leading to performance discrepancies. Therefore, it is important to know about the characteristics of the recording device; however, it is difficult to know the device's behavior in most cases. In this study, we propose a diffusion-model-based device characteristic transfer to estimate the device's frequency response only with the recorded signals. By joint-training the conditional and unconditional diffusion models, it is found that non-linear distortions and some filtered signals are reflected more than by only training the conditional model. We show that the proposed method transfers the style closely to the ground truth not only visually on the spectrogram but also the t-distributed stochastic neighbor embedding distribution and the performance of the device classifier. We also show the proposed method enhancing the performance as a data augmentation method for acoustic scene classification.

키워드

Acoustic Signal ProcessingAudio AcousticsAudio RecordingsAudio Signal ProcessingAudio SystemsClassification (of Information)DiffusionRecording InstrumentsAudio SignalDevice CharacteristicsDiffusion ModelModel-based OpcNon-linear DistortionsPerformanceRecorded SignalsRecording DevicesScene ClassificationSignal Processing SystemsStochastic SystemsAcoustic signal processingAudio acousticsAudio recordingsAudio signal processingAudio systemsClassification (of information)DiffusionRecording instruments
제목
Diffusion-based Target Device Style Transfer for Robust Acoustic Scene Classification
저자
Choi, Won-GookChang, Joon-Hyuk
DOI
10.1109/ICASSP49660.2025.10888162
발행일
2025-03
유형
Conference paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
1 ~ 5