메타학습을 이용한 시각-언어 모델 프롬프트 튜닝

Prompt Tuning for Vision-Language Models via Meta-Training

초록

Recently, applying large pre-trained vision-language models such as CLIP to various downstream tasks has shown good performance. In low-shot image classification, simply optimizing continuous prompt vectors emerged, but has the limitation of low generalizability. Recent research introduces additional structures and algorithms to solve this problem, but there is a disadvantage of inefficiency. So, we propose a meta-training framework utilizing multi-modal features to optimize the prompt vectors. Our proposed method achieves over a 9.6% accuracy improvement compared to other models when evaluated on both seen and unseen domains comprehensively, which has no any additional memory overhead or inference latency for zero-shot inference.

키워드

Vision-Language ModelsPrompt TuningMeta-LearningLow-shot Image Classification
제목
메타학습을 이용한 시각-언어 모델 프롬프트 튜닝
제목 (타언어)
Prompt Tuning for Vision-Language Models via Meta-Training
저자
김도현백성용
DOI
10.5909/JBE.2025.30.4.571
발행일
2025-07
유형
Y
저널명
방송공학회 논문지
30
4
페이지
571 ~ 579