Multimodal Emotion Recognition with Target Speaker-Based Facial Embeddings

Citations

SCOPUS

1

초록

Effectively recognizing emotions requires sophisticated approaches for interpreting diverse modalities, particularly in real-world scenarios where multiple data sources, such as speech, text, and visual cues, are often noisy and incomplete. This study proposes an advanced multimodal emotion recognition system that integrates these three modalities by adding the speaker detection and extraction algorithm within visual data. The pre-trained Q-Former used in the proposed system then captures and interprets visual signals supported with designated prompts, resulting in facial-related features that significantly improve emotion recognition performance. We then utilize a cross-modal transformer to unify the visual, speech, and text embeddings for accurate emotion classification. We achieved a 2.9% and 3.3% improvement in accuracy and F1 score, respectively, on the MELD dataset compared to the baseline.

키워드

cross-modal attentionmultimodal emotion recognitionquery transformertarget speaker
제목
Multimodal Emotion Recognition with Target Speaker-Based Facial Embeddings
저자
Heo, SerinKyung, JehyunChang, Joon-Hyuk
DOI
10.1109/ICASSP49660.2025.10888205
발행일
2025-03
유형
Conference paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
1 ~ 5