상세 보기
Integration of Global and Local Representations for Fine-Grained Cross-Modal Alignment
- Jin, Seungwan;
- Choi, Hoyoung;
- Noh, Taehyung;
- Han, Kyungsik
WEB OF SCIENCE
1SCOPUS
1초록
Fashion is one of the representative domains of fine-grained Vision-Language Pre-training (VLP) involving a large number of images and text. Previous fashion VLP research has proposed various pre-training tasks to account for fine-grained details in multimodal fusion. However, fashion VLP research has not yet addressed the need to focus on (1) uni-modal embeddings that reflect fine-grained features and (2) hard negative samples to improve the performance of fine-grained V+L retrieval tasks. In this paper, we propose Fashion-FINE (Fashion VLP with Fine-grained Cross-modal Alignment using the INtegrated representations of global and local patch Embeddings), which consists of three key modules. First, a modality-agnostic adapter (MAA) learns uni-modal integrated representations and reflects fine-grained details contained in local patches. Second, hard negative mining with focal loss (HNM-F) performs cross-modal alignment using the integrated representations, focusing on hard negatives to boost the learning of fine-grained cross-modal alignment. Third, comprehensive cross-modal alignment (C-CmA) extracts low- and high-level fashion information from the text and learns the semantic alignment to encourage disentangled embedding of the integrated image representations. Fashion-FINE achieved state-of-the-art performance on two representative public benchmarks (i.e., FashionGen and FashionIQ) in three representative V+L retrieval tasks, demonstrating its effectiveness in learning fine-grained features.
키워드
- 제목
- Integration of Global and Local Representations for Fine-Grained Cross-Modal Alignment
- 저자
- Jin, Seungwan; Choi, Hoyoung; Noh, Taehyung; Han, Kyungsik
- 발행일
- 2024-11
- 유형
- Proceedings Paper
- 권
- 15141
- 페이지
- 53 ~ 70