Audio Captioning Using Semantic Alignment Enhancer

Citations

SCOPUS

1

초록

Automated audio captioning (AAC) relies on the meaningful transfer of information from the encoder to the decoder. We propose a module called the semantic alignment enhancer to improve this process. This module converts the keywords extracted by the keyword extractor from audio features into captions. It then minimizes the distance between the text embeddings, output of the decoder, and the keyword caption embeddings, enhancing semantic similarity of alignment between audio features and text features. Additionally, to address the issue of data scarcity in audio captioning, we employ transfer learning. Specifically, we pre-train an audio encoder using the recently released large-scale audio-text paired WavCaps dataset. We use this pre-trained encoder to enhance the model's performance during the fine-tuning process. As a result, this approach enhances the alignment between audio features and textual information, which leads to improved performance in AAC tasks on the Clotho dataset compared to the baseline.

키워드

Audio CaptioningAudio TaggingSemantic AlignmentAudio captioningAudio featuresAudio taggingData scarcityEmbeddingsSemantic alignmentsSemantic similarityText featureTransfer learningTransfer of information
제목
Audio Captioning Using Semantic Alignment Enhancer
저자
박윤아Chang, Joon-Hyuk
DOI
10.1109/IC-NIDC59918.2023.10390585
발행일
2023-11
유형
Conference paper
저널명
Proceedings of 2023 8th IEEE International Conference on Network Intelligence and Digital Content, IC-NIDC 2023
페이지
374 ~ 378