CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation

  • Oh, Hyunwoo
  • Cha, Seung-ju
  • Lee, Kwanyoung
  • Kim, Si-woo
  • Kim, Dongjin
Citations

SCOPUS

2

초록

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining ) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector ). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.

키워드

audio to image generationdiffusion modellanguage-guided generationmulti-modal representationAlignmentClassification (of information)Image codingSignal encoding
제목
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
저자
Oh, HyunwooCha, Seung-juLee, KwanyoungKim, Si-wooKim, Dongjin
DOI
10.1145/3746027.3755130
발행일
2025-10
유형
Conference paper
저널명
MM 2025 - Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with MM 2025
페이지
9773 ~ 9782